✳flâneur — a map of the web's best reading
Nathan Chen
13 followers · 19 following · 1620 views
Open this reading profile →
on the atlas — 69
- Huxley Marvit1 savers
- Akira Yoshiyama4 savers
- How to Think About TPUs | How To Scale Your Model5 savers
- So You Want To Make Marginal Progress... — LessWrong2 savers
- Yudhister Kumar3 savers
- DeltaNet Explained (Part I) | Songlin Yang4 savers
- depression handbook | writing10 savers
- A short note on some aspects of long context attention | nor's blog1 savers
- Physics of Language Models - Part 3.3: Scaling Laws1 savers
- Linear Attention Fundamentals | Hailey Schoelkopf2 savers
- 1.5x faster MoE training with custom MXFP8 kernels · Cursor1 savers
- Best practices to accelerate inference for large-scale production workloads1 savers
- Lecture 21: Eigenvalues and eigenvectors1 savers
- five-thirty, again | writing1 savers
- Defeating the Training-Inference Mismatch via FP161 savers
- Kevin Wang1 savers
- errorgorn’s blog1 savers
- A Visual Guide to Mamba and State Space Models2 savers
- Specification gaming: the flip side of AI ingenuity - Google DeepMind3 savers
- Curius / Onboarding2507 savers
- Dario Amodei — Machines of Loving Grace50 savers
- escaping flatland: career advice for CS undergrads41 savers
- Becoming a magician – Autotranslucence39 savers
- No one can teach you to have conviction | benkuhn.net38 savers
- A Mathematical Framework for Transformer Circuits36 savers
- On the Biology of a Large Language Model31 savers
- How To Scale Your Model26 savers
- On-Policy Distillation - Thinking Machines Lab26 savers
- Making Deep Learning Go Faster25 savers
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklog23 savers
- Everything is Fertile16 savers
- Dario Amodei — On DeepSeek and Export Controls16 savers
- AGI Ruin: A List of Lethalities - LessWrong15 savers
- Deriving Muon15 savers
- Transformer Circuits Thread14 savers
- against narrative - by rayne fisher-quann14 savers
- The Smol Training Playbook: The Secrets to Building World-Class LLMs - a Hugging Face Space by HuggingFaceTB13 savers
- The Second Half – Shunyu Yao – 姚顺雨13 savers
- Illustrating Reinforcement Learning from Human Feedback (RLHF)13 savers
- We Induced Smells With Ultrasound11 savers
- mason wang11 savers
- Welcome to Spinning Up in Deep RL! — Spinning Up documentation11 savers
- A (Long) Peek into Reinforcement Learning | Lil'Log10 savers
- Part 1: Key Concepts in RL — Spinning Up documentation10 savers
- Growing Neural Cellular Automata9 savers
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordić9 savers
- All About Rooflines | How To Scale Your Model9 savers
- On the Tradeoffs of SSMs and Transformers | Goomba Lab8 savers
- Simon Willison’s Weblog8 savers
- Hopfield network7 savers
- Inside vLLM: Anatomy of a High-Throughput LLM Inference System - Aleksa Gordić7 savers
- Some thoughts on writing6 savers
- Q-learning is not yet scalable6 savers
- On idea-driven ideas5 savers
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Research5 savers
- Writing one sentence per line | Derek Sivers5 savers
- Quarter Mile4 savers
- ELI5: FlashAttention. Step by step explanation of how one of… | by Aleksa Gordić | Jul, 2023 | Medium4 savers
- Blog | karpathy4 savers
- Fermi Estimates - LessWrong4 savers
- Stop Climbing!4 savers
- Essays4 savers
- AI Control: Improving Safety Despite Intentional Subversion — LessWrong4 savers
- Kognise!4 savers
- the bug that taught me more about PyTorch than years of using it | Elana Simon3 savers
- Physical Intelligence (π)3 savers
- How to Write Usefully2 savers
- Post 50: Good Research Takes are Not Sufficient for Good Strategic Takes — Neel Nanda2 savers
- Pipeline-Parallelism: Distributed Training via Model Partitioning2 savers