Auditing Games for Sandbagging [paper] — LessWrong
Jordan Taylor, Sid Black, Dillon Bowen, Thomas Read, Satvik Golechha, Alex Zelenka-Martin, Oliver Makins, Connor Kissane, Kola Ayonrinde, Jacob Meriz…
x Auditing Games for Sandbagging [paper] — LessWrong AI Auditing AI Frontpage 2025 Top Fifty: 14 % 103 Auditing Games for Sandbagging [paper] by Jordan Taylor , Joseph Bloom 9th Dec 2025 AI Alignment Forum 12 min read 4 103 Ω 52 Jordan Taylor, Sid Black, Dillon Bowen, Thomas Read, Satvik Golechha, Alex Zelenka-Martin, Oliver Makins, Connor Kissane, Kola Ayonrinde, Jacob Merizian, Samuel Marks, Chris Cundy, Joseph Bloom UK AI Security Institute, FAR.AI , Anthropic Links: Paper | Code | Models | Transcripts | Interactive Demo Epistemic Status: We're sharing our paper and a hastily written summar
Explore this link on the map →related reading
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Sandbagging with misaligned action - Chain-of-Thought Transcript - Anti-Schemingantischeming.ai
- Automated Researchers Can Subtly Sandbagalignment.anthropic.com
- 2312.06942arxiv.org
- Auditing language models for hidden objectives — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- AuditBenchalignment.anthropic.com
- Building and evaluating alignment auditing agents — AI Alignment Forumalignmentforum.org