DeepSeek and the Day Before New Year's
By now, almost everyone in the AI/ML world knows of the bombshell that the Chinese company DeepSeek dropped in January 2025, with their release of DeepSeek-R1 and DeepSeek-R1-Zero large language models. Reactions ranged from cautious awe to downright skepticism, with some suggesting that R1 was simply a distilled version of an OpenAI frontier model. Such claims notwithstanding, the models incorporated architectural design decisions that were evidence of a research lab hard at work. Then, on 31 Dec, 2025, to cap off an already influential year (during which R1 reportedly became the first LLM to be peer-reviewed in a major journal), DeepSeek published mHC: Manifold-Constrained Hyper-Connections, a principled improvement on an existing deep neural network architecture. The Internet lit up again, reminiscent of the world's earlier encounter with DeepSeek. This time, however, the tone was distinctly different; many were appreciative of an advance that didn't involve simply throwing more com
Explore this link on the map →