✳flâneur — a map of the web's best reading
Models don’t seem to be dishonest in the way humans are — LessWrong
lesswrong.com · 3,014 words · saved by 1 readers
TLDR * Models often behave dishonestly without acquiring a coherent deceptive disposition. * We trained some mid-sized models on their own plausibl…
x Models don’t seem to be dishonest in the way humans are — LessWrong AI Frontpage 42 Models don’t seem to be dishonest in the way humans are by David Africa , Jacob Pfau 22nd Jul 2026 11 min read 2 42 TLDR Models often behave dishonestly without acquiring a coherent deceptive disposition. We trained some mid-sized models on their own plausible but false reasoning. True and false training usually produced nearly identical downstream effects. Even statements contradicting latent knowledge transferred only weakly to unrelated dishonesty. General deception may require agency, persistent private i
Explore this link on the map →related reading
- How confessions can keep language models honest | OpenAIopenai.com
- confessions_paper.pdfcdn.openai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Alignment faking in large language modelsarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Why we are excited about confession! — LessWronglesswrong.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org