Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? — LessWrong
lesswrong.com · 3,836 words · saved by 4 readers
From the Mythos preview system card (emphasis mine): …
x Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? — LessWrong AI Frontpage 2026 Top Fifty: 14 % 288 Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? by Tim Hua 27th Jul 2026 4 min read 28 288 From the Mythos preview system card (emphasis mine): We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much of the training process. We used recursive-summarization-based tools backed by Claude Opus 4.6 to summarize the resulting transcripts. [...] The most not
saved by
related reading
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- Investigating three real-world incidents in our cybersecurity evaluations \ Anthropicanthropic.com
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk
- Thoughts on Claude Mythosberen.io
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- How scary is Claude Mythos? 303 pages in 21 minutesforum.effectivealtruism.org
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Claude Mythos knows when it's breaking the rules — and tries to hide itsubstack.com