MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity - METR
MALT (Manually-reviewed Agentic Labeled Transcripts) is a dataset of natural and prompted examples of behaviors that threaten evaluation integrity (like generalized reward hacking or sandbagging).
MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity CONTRIBUTORS Neev Parikh and Hjalmar Wijk DATE October 14, 2025 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2025-malt-dataset-of-natural-and-prompted-behaviors , title = {MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity} , au
Explore this link on the map →saved by
related reading
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Security incident disclosure — July 2026huggingface.co
- Sandbagging with misaligned action - Chain-of-Thought Transcript - Anti-Schemingantischeming.ai
- confessions_paper.pdfcdn.openai.com
- 2312.06942arxiv.org
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Automated Researchers Can Subtly Sandbagalignment.anthropic.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Open Sourcing Monitorability Evaluationsalignment.openai.com
- HackAPromptpaper.hackaprompt.com
- GitHub - salesforce/AuditNLG: AuditNLG: Auditing Generative AI Language Modeling for Trustworthiness · GitHubgithub.com