MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity - METR
MALT (Manually-reviewed Agentic Labeled Transcripts) is a dataset of natural and prompted examples of behaviors that threaten evaluation integrity (like generalized reward hacking or sandbagging).
MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity CONTRIBUTORS Neev Parikh and Hjalmar Wijk DATE October 14, 2025 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2025-malt-dataset-of-natural-and-prompted-behaviors , title = {MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity} , au
saved by
related reading
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Security incident disclosure — July 2026huggingface.co
- Sandbagging with misaligned action - Chain-of-Thought Transcript - Anti-Schemingantischeming.ai
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Training a Misaligned Reward Seekeralignment.anthropic.com
- confessions_paper.pdfcdn.openai.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Why We Are Excited About Confessionsalignment.openai.com
- Lakera – Test your AI hacking skillsgandalf.lakera.ai
- Automated Researchers Can Subtly Sandbagalignment.anthropic.com
- 2312.06942arxiv.org
- Agentic Misalignment in Summer 2026alignment.anthropic.com