DeepSeek and the Day Before New Year's
By now, almost everyone in the AI/ML world knows of the bombshell that the Chinese company DeepSeek dropped in January 2025, with their release of DeepSeek-R1 and DeepSeek-R1-Zero large language models. Reactions ranged from cautious awe to downright skepticism, with some suggesting that R1 was simply a distilled version of an OpenAI frontier model. Such claims notwithstanding, the models incorporated architectural design decisions that were evidence of a research lab hard at work. Then, on 31 Dec, 2025, to cap off an already influential year (during which R1 reportedly became the first LLM to be peer-reviewed in a major journal), DeepSeek published mHC: Manifold-Constrained Hyper-Connections, a principled improvement on an existing deep neural network architecture. The Internet lit up again, reminiscent of the world's earlier encounter with DeepSeek. This time, however, the tone was distinctly different; many were appreciative of an advance that didn't involve simply throwing more com
By now, almost everyone in the AI/ML world knows of the bombshell that the Chinese company DeepSeek dropped in January 2025, with their release of DeepSeek-R1 and DeepSeek-R1-Zero large language models. Reactions ranged from cautious awe to downright skepticism, with some suggesting that R1 was simply a distilled version of an OpenAI frontier model. Such claims notwithstanding, the models incorporated architectural design decisions that were evidence of a research lab hard at work. Then, on 31 Dec, 2025, to cap off an already influential year (during which R1 reportedly became the first LLM…
saved by
related reading
- DeepSeek: The View from Chinachinatalk.media
- Deepseek: The Quiet Giant Leading China’s AI Racechinatalk.media
- Dario Amodei — On DeepSeek and Export Controlsdarioamodei.com
- DeepSeek-R1arxiv.org
- WIRTW: Deepseek Breakthroughchamath.substack.com
- Alfred Lin (@Alfred_Lin) on Xx.com
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- arxiv.org/pdf/2512.24880#page=3.56arxiv.org
- The Decade of Deep Learning | Leo Gaobmk.sh
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- DeepSeek-R1 and exploring DeepSeek-R1-Distill-Llama-8Bsimonwillison.net
- mHC: Manifold-Constrained Hyper-Connectionsalphaxiv.org