Arya G
1 followers · 167 views
on the atlas — 34
- Inside vLLM: Anatomy of a High-Throughput LLM Inference System - Aleksa Gordić11 savers
- FLT: Anthropic has beaten me to it | Xena4 savers
- Formalizing Fermat's Last Theorem \ Anthropic5 savers
- How Claude Performs on Robotics Tasks \ Anthropic6 savers
- How our data shaped neural architecture discovery, and how automation can reshape the future | Core Automation5 savers
- Erdős Problems4 savers
- Announcing FrontierMath Erdős | Epoch AI1 savers
- [2607.16051] Loop the Loopies!1 savers
- Harness Engineering for Self-Improvement | Lil'Log19 savers
- DeepSeek-V3: A Large-Scale MoE Pretraining Benchmark for MLPerf Training v6.0 - MLCommons1 savers
- Frontier-scale RL with Kimi K3 | Applied Compute1 savers
- Defeating Nondeterminism in LLM Inference - Thinking Machines Lab40 savers
- Non-determinism in GPT-4 is caused by Sparse MoE - 152334H5 savers
- The 4-bitter Lesson | humans&5 savers
- [2512.24880] mHC: Manifold-Constrained Hyper-Connections1 savers
- [Hero Run] 535B-A23B on 18T tokens · Issue #8435 · marin-community/marin1 savers
- Mixture-of-Kittens: our open-source MoE megakernel for NVL72s · Cursor3 savers
- Better MoE model inference with warp decode · Cursor3 savers
- Mixture of Experts Quantile Balancing: Validated at 32B-A5B (1e22 FLOPs) Scale | Open Athena1 savers
- Open Athena | Scaling Laws That Extrapolate 300× Past the Fit2 savers
- [2604.20920] Simplified Sparse Attention via Gist Tokens1 savers
- [2108.12409] Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation2 savers
- Four LLM loss functions → four flavors of LLM misalignment — LessWrong3 savers
- The Curious Case of the bos_token — LessWrong1 savers
- How GPT-5, Claude, and Gemini are actually trained and served – Reiner Pope - YouTube2 savers
- [2509.14786] Pre-training under infinite compute2 savers
- [2608.09867] Stealing Reasoning Traces from Proprietary LLM APIs2 savers
- llms-cant-jump.pdf1 savers
- Hot Take: LLM can 'jump' | Yong Zheng-Xin1 savers
- The Scaling Hypothesis · Gwern.net16 savers
- Scaling Laws, Carefully | Lil'Log15 savers
- A retrospective of AI alignment14 savers
- Your Agents Are Not Time Aware — LessWrong3 savers
- A short note on some aspects of long context attention | nor's blog2 savers