What I learned this week - Pretraining parallelisms, Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RL, On why pretraining runs fails
dwarkesh.com · 1,305 words · saved by 4 readers
April 15, 2025
Blog What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RL April 15, 2025 Dwarkesh Patel Apr 15, 2026 133 16 6 Share At the end of my conversation with Michael Nielsen , we talked about how to actually retain what you learn. Michael’s advice was to make some kind of demanding artifact. Write something up. Try to explain it. So in that spirit, here are notes on some topics I’ve learned about over the last week or two. These notes are extremely rough, and have many mistakes. Can distillation be stopped? Can the frontier labs stop distillati
saved by
related reading
- AI in 2025: gestalt — LessWronglesswrong.com
- Stealing Reasoning Traces from Proprietary LLM APIsresearch.snyk.io
- Thoughts on AI progress (Dec 2025)substack.com
- Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? — LessWronglesswrong.com
- Detecting and preventing distillation attacks \ Anthropicanthropic.com
- Research — Paradigm 3paradigm3.org
- RL Environments and RL for Science: Data Foundries and Multi-Agent Architecturesnewsletter.semianalysis.com
- AI 2027ai-2027.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- My picture of the present in AI — LessWronglesswrong.com
- Composer2.pdfcursor.com