Do your capabilities homework — LessWrong
It seems to me that a lot of technical ai safety people haven't done their capabilities homework - and that's a shame! I'll try to illuminate here ma…
x Do your capabilities homework — LessWrong AI Frontpage 57 Do your capabilities homework by RobinHa 1st Aug 2026 6 min read 11 57 It seems to me that a lot of technical ai safety people haven't done their capabilities homework - and that's a shame! I'll try to illuminate here mainly with an example as to why I think people who care about safety should totally pay more attention to the trends and actively engage with them - the case for safe AI not through an additional loss term but as a consequence of the learning algorithm! RLVR It's now been 1.5 years since R1 came out - the paper which re
Explore this link on the map →related reading
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsnoahziems.com
- AI in 2025: gestalt — LessWronglesswrong.com
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- ar0cket1 on X: "Solving OPSD (basically)" / Xx.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- AI safety techniques leveraging distillation — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- GRPO is terrible — LessWronglesswrong.com