✳flâneur — a map of the web's best reading
Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? — LessWrong
lesswrong.com · 3,836 words · saved by 1 readers
From the Mythos preview system card (emphasis mine): …
x Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? — LessWrong AI Frontpage 2026 Top Fifty: 14 % 288 Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? by Tim Hua 27th Jul 2026 4 min read 28 288 From the Mythos preview system card (emphasis mine): We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much of the training process. We used recursive-summarization-based tools backed by Claude Opus 4.6 to summarize the resulting transcripts. [...] The most not
Explore this link on the map →related reading
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Thoughts on Claude Mythosberen.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- My picture of the present in AI — LessWronglesswrong.com
- Claude Fable 5 and Claude Mythos 5 \ Anthropicanthropic.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Cybersecurity Looks Like Proof of Work Nowdbreunig.com
- Teaching Claude why \ Anthropicanthropic.com