flâneur — a map of the web's best reading

Thoughts on the conservative assumptions in AI control

blog.redwoodresearch.org · 3,819 words · saved by 2 readers

Why are we so friendly to the red team?

Thoughts on the conservative assumptions in AI control Why are we so friendly to the red team? Buck Shlegeris Jan 17, 2025 5 Share Work that I’ve done on techniques for mitigating risk from misaligned AI often makes a number of conservative assumptions about the capabilities of the AIs we’re trying to control. (E.g. the original AI control paper , Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats , How to prevent collusion when using untrusted models to monitor each other .) For example: The AIs are consistently trying to subvert safety measures. They’re very good at strategizi

Explore this link on the map →

saved by

related reading