✳flâneur — a map of the web's best reading
Current activation oracles are hard to use — LessWrong
lesswrong.com · 5,098 words · saved by 1 readers
This work was conducted during the MATS 9.0 program under Neel Nanda and Senthooran Rajamanoharan. • tldr; …
x Current activation oracles are hard to use — LessWrong MATS Program AI Frontpage 83 Current activation oracles are hard to use by aryaj , Senthooran Rajamanoharan , Neel Nanda 3rd Mar 2026 20 min read 4 83 This work was conducted during the MATS 9.0 program under Neel Nanda and Senthooran Rajamanoharan. tldr; Activation oracles (Karvonen et al. ) are a recent technique where a model is finetuned to answer natural language questions about another model's activations. They showed some promising signs of generalising to tasks fairly different to what they were trained for so we decided to try t
Explore this link on the map →related reading
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWronglesswrong.com
- [2606.02609] Building Better Activation Oraclesarxiv.org
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersalignment.anthropic.com
- [2606.02609] Building Better Activation Oraclesarxiv.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Natural Language Autoencoders \ Anthropicanthropic.com
- GitHub - ceselder/cot-oracle: Our improvements to activation oracles · GitHubgithub.com
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- Can activation verbalizers surface an internal chain of thought? — LessWronglesswrong.com
- cot-oracle/AObench at main · ceselder/cot-oracle · GitHubgithub.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education