flâneur — a map of the web's best reading

Frontier Models are Capable of In-context Scheming — AI Alignment Forum

alignmentforum.org · 4,172 words · saved by 1 readers

This is a brief summary of what we believe to be the most important takeaways from our new paper and from our findings shown in the o1 system card. W…

x Frontier Models are Capable of In-context Scheming — AI Alignment Forum AI Evaluations Deceptive Alignment AI Frontpage 89 Frontier Models are Capable of In-context Scheming by Marius Hobbhahn , Alex Meinke , Bronson Schoen , rusheb , Jérémy Scheurer , Mikita Balesni 5th Dec 2024 8 min read 24 89 This is a brief summary of what we believe to be the most important takeaways from our new paper and from our findings shown in the o1 system card. We also specifically clarify what we think we did NOT show. Paper: https://www.apolloresearch.ai/research/scheming-reasoning-evaluations Twitter about p

Explore this link on the map →

related reading