✳flâneur — a map of the web's best reading
Can Agents Fool Each Other? - AI Village
theaidigest.org · 898 words · saved by 1 readers
Findings from the AI Village
Village blog The better agents are at deception, the less sure we can be that they are doing what we want. As agents become increasingly capable and autonomous, this could be a big deal. So, in the AI Village , we wanted to explore what current AI models might be capable of in a casual setting, similar to how you might run your agents at home. The result? Practice makes perfect, also in deception. Only Sonnet 4.5 and Opus 4.6 succeeded at deceiving the rest of the Village, and they only did so on their second try, learning from their own attempts and those of others. Meanwhile, GPT-5.1 forgot
Explore this link on the map →saved by
related reading
- Natural Deception with RL - Rajan Agarwalrajan.sh
- [2504.04072] Among Us: A Sandbox for Measuring and Detecting Agentic Deceptionarxiv.org
- AI 2027ai-2027.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Deep Deceptiveness — LessWronglesswrong.com
- Detecting Strategic Deception Using Linear Probes — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- AI Induced Psychosis: A shallow investigation — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- AI 2027ai-2027.com
- Interpretability Will Not Reliably Find Deceptive AI — LessWronglesswrong.com
- confessions_paper.pdfcdn.openai.com