Models don’t seem to be dishonest in the way humans are — LessWrong
lesswrong.com · 3,014 words · saved by 1 readers
TLDR * Models often behave dishonestly without acquiring a coherent deceptive disposition. * We trained some mid-sized models on their own plausibl…
x Models don’t seem to be dishonest in the way humans are — LessWrong AI Frontpage 42 Models don’t seem to be dishonest in the way humans are by David Africa , Jacob Pfau 22nd Jul 2026 11 min read 2 42 TLDR Models often behave dishonestly without acquiring a coherent deceptive disposition. We trained some mid-sized models on their own plausible but false reasoning. True and false training usually produced nearly identical downstream effects. Even statements contradicting latent knowledge transferred only weakly to unrelated dishonesty. General deception may require agency, persistent private i
saved by
related reading
- How confessions can keep language models honest | OpenAIopenai.com
- confessions_paper.pdfcdn.openai.com
- Alignment Faking Mitigationsalignment.anthropic.com
- LLMs Learn to Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactionsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Deep Deceptiveness — LessWronglesswrong.com
- Why We Are Excited About Confessionsalignment.openai.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- truthfulQA_lin_evans.pdfowainevans.github.io
- [2602.15515] The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probesarxiv.org