flâneur — a map of the web's best reading

Control protocols don’t always need to know which models are scheming — LessWrong

lesswrong.com · 1,833 words · saved by 1 readers

These are my personal views. • To detect if an agent is taking a catastrophically dangerous action, you might want to monitor its actions using the s…

x Control protocols don’t always need to know which models are scheming — LessWrong AI Frontpage 39 Control protocols don’t always need to know which models are scheming by Fabien Roger 26th Apr 2026 7 min read 1 39 These are my personal views. To detect if an agent is taking a catastrophically dangerous action, you might want to monitor its actions using the smartest model that is too weak to be a schemer. But knowing what models are weak enough that they are unlikely to scheme is difficult, which puts you in a difficult spot: take a model too strong and it might actually be a schemer and lie

Explore this link on the map →

saved by

related reading