flâneur — a map of the web's best reading

Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs

blog.redwoodresearch.org · 8,715 words · saved by 1 readers

When a misaligned AI can only output tiny amounts of information, it may find sabotage very difficult

Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs When a misaligned AI can only output tiny amounts of information, it may find sabotage very difficult Caleb Biddulph and Adam Kaufman Jul 27, 2026 2 1 Share TL;DR: We introduce the untrusted advice protocol, in which a trusted executor LLM takes every action and an untrusted advisor LLM can only send it short hints. Even with as few as 4 characters per step, this advice recovers a substantial fraction of the capability gap between the two models. Because the untrusted LLM’s influence flows through such a narr

Explore this link on the map →

related reading