flâneur — a map of the web's best reading

Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? — LessWrong

lesswrong.com · 3,836 words · saved by 1 readers

From the Mythos preview system card (emphasis mine): …

x Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? — LessWrong AI Frontpage 2026 Top Fifty: 14 % 288 Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? by Tim Hua 27th Jul 2026 4 min read 28 288 From the Mythos preview system card (emphasis mine): We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much of the training process. We used recursive-summarization-based tools backed by Claude Opus 4.6 to summarize the resulting transcripts. [...] The most not

Explore this link on the map →

related reading