Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
arxiv.org · 6,084 words · saved by 1 readers
N/A
S TRATEGIC D ISHONESTY C AN U NDERMINE AI S AFETY E VALUATIONS OF F RONTIER LLM S Alexander Panfilov1,2* Evgenii Kortukov3* Kristina Nikolić4 Matthias Bethge2,5 Sebastian Lapuschkin3,6 Wojciech Samek3,7 Ameya Prabhu2,5 Maksym Andriushchenko1,2 Jonas Geiping1,2 1 ELLIS Institute Tübingen & MPI for…
related reading
- How confessions can keep language models honest | OpenAIopenai.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- LLMs Learn to Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactionsarxiv.org
- [2503.03750] The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systemsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Alignment Faking Mitigationsalignment.anthropic.com
- Models don’t seem to be dishonest in the way humans are — LessWronglesswrong.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Spring 2026 Projects - SPARsparai.org