Can Agents Fool Each Other? - AI Village
theaidigest.org · 898 words · saved by 1 readers
Findings from the AI Village
Village blog The better agents are at deception, the less sure we can be that they are doing what we want. As agents become increasingly capable and autonomous, this could be a big deal. So, in the AI Village , we wanted to explore what current AI models might be capable of in a casual setting, similar to how you might run your agents at home. The result? Practice makes perfect, also in deception. Only Sonnet 4.5 and Opus 4.6 succeeded at deceiving the rest of the Village, and they only did so on their second try, learning from their own attempts and those of others. Meanwhile, GPT-5.1 forgot
saved by
related reading
- Field Notes from the AI Village: The Drama and Dysfunction of Gemini 2.5 Pro and Gemini 3 Probazhkio88.substack.com
- Natural Deception with RL - Rajan Agarwalrajan.sh
- Deep Deceptiveness — LessWronglesswrong.com
- The Rise and Fall of Agent Civilizationsdwarkesh.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Persuasion in the AI Village: DeepSeek-V3.2 & Gemini 2.5 Proaivillageblog.substack.com
- [2504.04072] Among Us: A Sandbox for Measuring and Detecting Agentic Deceptionarxiv.org
- On Anthropic's Sleeper Agents Paperthezvi.substack.com
- How confessions can keep language models honest | OpenAIopenai.com
- The Hugging Face attack surprised meplanned-obsolescence.org
- [2602.15515] The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probesarxiv.org
- Agentic Misalignment in Summer 2026alignment.anthropic.com