flâneur — a map of the web's best reading

Models don’t seem to be dishonest in the way humans are — LessWrong

lesswrong.com · 3,014 words · saved by 1 readers

TLDR * Models often behave dishonestly without acquiring a coherent deceptive disposition. * We trained some mid-sized models on their own plausibl…

x Models don’t seem to be dishonest in the way humans are — LessWrong AI Frontpage 42 Models don’t seem to be dishonest in the way humans are by David Africa , Jacob Pfau 22nd Jul 2026 11 min read 2 42 TLDR Models often behave dishonestly without acquiring a coherent deceptive disposition. We trained some mid-sized models on their own plausible but false reasoning. True and false training usually produced nearly identical downstream effects. Even statements contradicting latent knowledge transferred only weakly to unrelated dishonesty. General deception may require agency, persistent private i

Explore this link on the map →

related reading