flâneur — a map of the web's best reading

Policy Options for Preserving Chain of Thought Monitorability — Institute for AI Policy and Strategy

iaps.ai · 432 words · saved by 1 readers

By using this website, you agree to our use of cookies. We use cookies to provide you with a great experience and to help our website run effectively. This report was co-authored by Oscar Delaney, Oliver Guest, and Renan Araujo. Today’s most advanced artificial intelligence (AI) models use “chain of thought” (CoT) reasoning; monitoring this CoT can be valuable for controlling these systems and ensuring that they will behave as intended. This is because, like humans, AIs need to think in steps in order to solve complex problems. We can monitor the steps that the model writes (i.e., the CoT) and intervene to stop the model’s default course of action, if a concerning intention is expressed in the CoT. As AI systems are deployed in higher stakes contexts, CoT monitoring could become increasingly important for preventing risks to public safety and national security. In experiments, researchers have already found AIs to cheat on a software development task by altering a test to make the exis

Explore this link on the map →

saved by