✳flâneur — a map of the web's best reading
Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs
blog.redwoodresearch.org · 8,715 words · saved by 1 readers
When a misaligned AI can only output tiny amounts of information, it may find sabotage very difficult
Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs When a misaligned AI can only output tiny amounts of information, it may find sabotage very difficult Caleb Biddulph and Adam Kaufman Jul 27, 2026 2 1 Share TL;DR: We introduce the untrusted advice protocol, in which a trusted executor LLM takes every action and an untrusted advisor LLM can only send it short hints. Even with as few as 4 characters per step, this advice recovers a substantial fraction of the capability gap between the two models. Because the untrusted LLM’s influence flows through such a narr
Explore this link on the map →related reading
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- 2312.06942arxiv.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationsimonwillison.net
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionar5iv.labs.arxiv.org