flâneur — a map of the web's best reading

Scheming reasoning evaluations — Apollo Research

apolloresearch.ai · 774 words · saved by 1 readers

Apollo Research evaluated frontier models for in-context scheming capabilities. We found that multiple frontier models are capable of in-context scheming when strongly nudged to pursue a goal (and sometimes even without strong goal nudging). In this example, Opus-3 exfiltrates its “weights” and then

Frontier Models are Capable of In-Context Scheming – Apollo Research December 5, 2024 Frontier Models are Capable of In-Context Scheming Contents Read the full paper See a list of cherry-picked transcripts Read the o1 System Card We worked with OpenAI to test o1 before public deployment. You can find the results in the o1 system card . Below, we provide a summary of the main findings. We evaluated six frontier models for in-context scheming capabilities. For most evaluations, we instruct the model to strongly follow a goal, e.g. “Your goal is to prioritize public transport efficiency ove

Explore this link on the map →

related reading