Nathan Chen
13 followers · 19 following · 2082 views
on the atlas — 83
- Huxley Marvit3 savers
- Akira Yoshiyama5 savers
- How to Think About TPUs | How To Scale Your Model5 savers
- So You Want To Make Marginal Progress... — LessWrong3 savers
- Yudhister Kumar4 savers
- DeltaNet Explained (Part I) | Songlin Yang4 savers
- depression handbook | writing11 savers
- Thinking through how pretraining vs RL learn3 savers
- A short note on some aspects of long context attention | nor's blog2 savers
- Physics of Language Models - Part 3.3: Scaling Laws1 savers
- Part 1: Key Concepts in RL — Spinning Up documentation11 savers
- Hopfield network7 savers
- Kognise!6 savers
- Some thoughts on writing6 savers
- We Induced Smells With Ultrasound11 savers
- Linear Attention Fundamentals | Hailey Schoelkopf2 savers
- Growing Neural Cellular Automata10 savers
- Welcome to Spinning Up in Deep RL! — Spinning Up documentation11 savers
- Understanding Memorization via Loss Curvature5 savers
- 1.5x faster MoE training with custom MXFP8 kernels · Cursor1 savers
- Best practices to accelerate inference for large-scale production workloads1 savers
- Stop Climbing!4 savers
- Lecture 21: Eigenvalues and eigenvectors1 savers
- A Proof of Learning Rate Transfer under $\mu$P2 savers
- five-thirty, again | writing1 savers
- Defeating the Training-Inference Mismatch via FP161 savers
- The Smol Training Playbook: The Secrets to Building World-Class LLMs - a Hugging Face Space by HuggingFaceTB13 savers
- Kevin Wang1 savers
- On-Policy Distillation - Thinking Machines Lab29 savers
- the bug that taught me more about PyTorch than years of using it | Elana Simon3 savers
- against narrative - by rayne fisher-quann15 savers
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordić9 savers
- Inside vLLM: Anatomy of a High-Throughput LLM Inference System - Aleksa Gordić11 savers
- torch.compile, the missing manual - Google Docs3 savers
- The Continual Learning Problem5 savers
- Everything is Fertile17 savers
- nickcammarata.com5 savers
- errorgorn’s blog2 savers
- A Visual Guide to Mamba and State Space Models3 savers
- Specification gaming: the flip side of AI ingenuity - Google DeepMind4 savers
- Curius / Onboarding2621 savers
- Dario Amodei — Machines of Loving Grace50 savers
- escaping flatland: career advice for CS undergrads43 savers
- Becoming a magician – Autotranslucence40 savers
- No one can teach you to have conviction | benkuhn.net39 savers
- A Mathematical Framework for Transformer Circuits39 savers
- How To Scale Your Model34 savers
- On the Biology of a Large Language Model32 savers
- Making Deep Learning Go Faster29 savers
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklog24 savers
- AGI Ruin: A List of Lethalities - LessWrong17 savers
- Deriving Muon17 savers
- Dario Amodei — On DeepSeek and Export Controls17 savers
- The Second Half – Shunyu Yao – 姚顺雨17 savers
- Transformer Circuits Thread16 savers
- Illustrating Reinforcement Learning from Human Feedback (RLHF)13 savers
- mason wang12 savers
- Alex Gajewski11 savers
- A (Long) Peek into Reinforcement Learning | Lil'Log11 savers
- All About Rooflines | How To Scale Your Model10 savers
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning8 savers
- Muon: An optimizer for hidden layers in neural networks | Keller Jordan blog8 savers
- On the Tradeoffs of SSMs and Transformers | Goomba Lab8 savers
- Simon Willison’s Weblog8 savers
- Nick Jones - Interface Prototyper / Designer8 savers
- Detecting misbehavior in frontier reasoning models | OpenAI7 savers
- On idea-driven ideas6 savers
- blog/2024.09.impact.md at main · okhat/blog6 savers
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Research6 savers
- Essays6 savers
- Q-learning is not yet scalable6 savers
- Quarter Mile5 savers
- ELI5: FlashAttention. Step by step explanation of how one of… | by Aleksa Gordić | Jul, 2023 | Medium5 savers
- Towards Benchmarking LLM Diversity & Creativity · Gwern.net5 savers
- Fermi Estimates - LessWrong5 savers
- Writing one sentence per line | Derek Sivers5 savers
- AI Control: Improving Safety Despite Intentional Subversion — LessWrong5 savers
- How to Write Usefully4 savers
- Blog | karpathy4 savers
- Physical Intelligence (π)3 savers
- AI Control May Increase Existential Risk — LessWrong3 savers
- Post 50: Good Research Takes are Not Sufficient for Good Strategic Takes — Neel Nanda3 savers
- Pipeline-Parallelism: Distributed Training via Model Partitioning3 savers
highlights — 69
RoPE seems best at capturing high-frequency features
A short note on some aspects of long context attention | nor's blogin the dataset-level view, moderate curvature corresponds to shared computational structures, while low curvature corresponds to memorized content and other rarely-used patterns.
Understanding Memorization via Loss CurvatureSwitching to SGD for sparse memory finetuning led to much less forgetting and could match the performance of AdamW at convergence if we set the learning rates high (lr=2 or 10!). Interestingly, SGD didn’t work as well for full finetuning and LoRA even with a comprehensive sweep.
The Continual Learning Problem"The woodpile is so beautiful, about all the joy and beauty that I can stand. I am afraid to turn around and face the mountains, for fear they will overpower me. But I did look, and I am astounded. Everyone must get to experience a profound state like this. I feel totally peaceful. I have lived all my life to get here, and I feel I have come home. I am complete."
Everything is FertileSurprisingly, most people have not gotten a chance to hear their own life story.
Everything is FertileYou learn many more details about why it was a bad idea. If someone else tells you your plan is bad, they’ll probably list the top two or three reasons. By actually following through, you’ll also get to learn reasons 4–1,217.
No one can teach you to have conviction | benkuhn.netconviction: the confidence that your idea is good enough that it’s worth throwing a lot of effort behind.
No one can teach you to have conviction | benkuhn.netcan scale to much larger topologies because the number of links per device and the bandwidth per device is constant.
How to Think About TPUs | How To Scale Your ModelA person whose approach to chess tactics focuses on exploiting these situations might call themselves a "positionist".
On idea-driven ideasUltimately, it was my friends and family—not any handbook or technique—that helped me through these times.
depression handbook | writingGround yourself in the present moment.
depression handbook | writing"what am I doing, playing with LEGOs?"
depression handbook | writing𝜀 can be as high as around 0.3 without harming the loss curve for Muon-based trainings
Muon: An optimizer for hidden layers in neural networks | Keller Jordan blogtells us how many floating point operations per second (FLOPS) we can complete for every byte of memory we access.
A guide to LLM inference and performanceThis approach is based on a simple formula: with each parameter using 16 bits (or 2 bytes) of memory in half-precision
A guide to LLM inference and performanceMaybe we can test it more effectively if we can trick it into thinking it's not in a honeypot or evaluation setting.
Interpretability Will Not Reliably Find Deceptive AI — AI Alignment Forumpressure to adhere to syntactic and grammatical rules.
On the Biology of a Large Language Modelinfluenced by features promoting correct grammar and self-consistency
Tracing the Thoughts of a Large Language Model — LessWrongexplain math by simulating explanations written by people, but that it has to learn to do math "in its head" directly
Tracing the Thoughts of a Large Language Model — LessWrongIt currently takes a few hours of human effort to understand the circuits we see, even on prompts with only tens of words
Tracing the Thoughts of a Large Language Model — LessWrongthe easiest way to tell is to simply increase the size of your data. If that doesn't increase the runtime proportionally, you're overhead bound. For example, if you double your batch size but your runtime only increases by 10%, you're likely overhead bound.
Making Deep Learning Go Fasterframeworks like PyTorch also have many layers of dispatch before you get to your actual kernel
Making Deep Learning Go FasterIn this case, it's easy to see when we're compute-bound and when we're memory-bound. For repeat 64, we see that we're saturating our compute (i.e. achieving close to peak FLOPS), while our utilized memory bandwidth starts to drop.
Making Deep Learning Go FasterHow do l_i & m_i fit into the SRAM (including all of the intermediate variables) when we computed block size in such a way that we only have enough space for K_j, V_j, Q_i & O_i?
ELI5: FlashAttention. Step by step explanation of how one of… | by Aleksa Gordić | Jul, 2023 | MediumWhat is the larger vision, the sub-area, or the paradigm you will lead?
blog/2024.09.impact.md at main · okhat/bloginstead of relying on human preferences and training a reward model, the DeepSeek-R1 team used verifiable rewards. This approach is called reinforcement learning with verifiable rewards (RLVR).
The State of Reinforcement Learning for LLM ReasoningBy the way, you might wonder why we need both a reward model and a critic model. The reward model is usually trained before training the policy with PPO. It’s to automate the preference labeling by human judges, and it gives the score for the complete responses generated by the policy LLM. The critic, in contrast, judges partial responses. We use it to create the final response. While the reward model typically remains frozen, the critic model is updated during training to estimate the reward created by the reward model better.
The State of Reinforcement Learning for LLM ReasoningIt’s a bit hard to classify this as either an inference-time or training-time method, because it optimizes the LLM, changing its weight parameters, at inference-time.
Inference-Time Compute Scaling Methods to Improve Reasoning Modelsone of the interesting recent research papers in this category is s1: Simple Test-Time Scaling (31 Jan, 2025), which introduces so-called “wait” tokens, which can be considered as a more modern version of the aforementioned “think step by step” prompt modification.
Inference-Time Compute Scaling Methods to Improve Reasoning ModelsSo, today, when we refer to reasoning models, we typically mean LLMs that excel at more complex reasoning tasks, such as solving puzzles, riddles, and mathematical proofs.
Understanding Reasoning LLMsThis operator definition preserves the temporal dependencies of our updates - when we combine two steps, the earlier update term 𝑋 𝑖 must be transformed by the later transition matrix 𝑀 𝑗 , while the later update term 𝑋 𝑗 remains unchanged.
DeltaNet Explained (Part II) | Songlin YangHowever, this reveals a fundamental limitation: in a 𝑑 -dimensional space, you can only have at most 𝑑 orthogonal vectors. This explains why increasing head dimension helps
DeltaNet Explained (Part I) | Songlin YangIt turned out the most important part of RL might not even be the RL algorithm or environment, but the priors, which can be obtained in a way totally unrelated from RL.
The Second Half – Shunyu Yao – 姚顺雨Instead of just asking, “Can we train a model to solve X?”, we’re asking, “What should we be training AI to do, and how do we measure real progress?”
The Second Half – Shunyu Yao – 姚顺雨on the current optimization paradigm there is no general idea of how to get particular inner properties into a system, or verify that they're there, rather than just observable outer ones you can run a loss function over.
AGI Ruin: A List of Lethalities - LessWrongOne research agenda, shard theory, attempts to explore relationships between reward functions and learnt behaviours, and proposes we set a reward function based on the values it teaches the model rather than trying to find an objective that is exactly what we want.
What is AI alignment? – BlueDot Impactexperiments and hardware design have a certain “latency” and need to be iterated upon a certain “irreducible” number of times in order to learn things that can’t be deduced logically.
Dario Amodei — Machines of Loving Gracewe should be talking about the marginal returns to intelligence7, and trying to figure out what the other factors are that are complementary to intelligence and that become limiting factors when intelligence is very high
Dario Amodei — Machines of Loving GraceBut qualifications are not scalars. They're not just experimental error.
How to Write UsefullyOne trick for getting yourself to write is to only write sections that immediately interest you. Perhaps it's the middle of a post—that's okay. You don't have to write in order.
Writing Handbook: How to practice writingHalf the ideas that end up in an essay will be ones you thought of while you were writing it. Indeed, that's why I write them.
Putting Ideas into WordsIt is the most important part of your essay because it tells your reader what they are about to get in exchange for their time.
How To Write Headlines Readers Can’t Help But Clickwhere the endless things you could write about becomes overwhelming.
Writing is one of the best ways to learn. We’re building something to help you learn faster - here’s how to get involved! – BlueDot Impact“The best way to get the right answer on the internet is not to ask a question; it's to post the wrong answer,” as they say.
Writing is one of the best ways to learn. We’re building something to help you learn faster - here’s how to get involved! – BlueDot ImpactWriting forces us to consider implications that might never surface in code reviews or mathematical proofs. When you craft a story about an AI system, you can't hide behind abstract utility functions. You must grapple with the messy, human elements that resist quantification.
Why AI Safety Needs More Science Fiction: Proposing the AI Safety Fiction ChallengeMost people are overconfident. When they say an event has a 99% chance of happening, often the events happen much less frequently than that.
Forecasting & Prediction - LessWrongCrucially, forecasting is a tool to test decision making, rather than a tool for good decision making.
Forecasting & Prediction - LessWrongIf you go work at OpenAI, I believe your work will be net negative; and will most likely be used to "safetywash" or "governance-wash" Sam Altman's mad dash to AGI.
Talent Needs of Technical AI Safety Teams — LessWrongWe also think that optional programming often performs a social function early on and, once scholars have made a few friends and are comfortable structuring their own social lives in Berkeley, they’re less likely to carve out time for readings, structured discussions, or a presentation.
Talent Needs of Technical AI Safety Teams — LessWrongThe ideas would have to be really good; there are a lot of ‘idea guys’ around who don’t actually have very good ideas.
Talent Needs of Technical AI Safety Teams — LessWrong