flâneur — a map of the web's best reading

Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWrong

lesswrong.com · 4,154 words · saved by 1 readers

TL;DR: We train LLMs to accept LLM neural activations as inputs and answer arbitrary questions about them in natural language. These Activation Oracl…

x Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWrong Interpretability (ML & AI) Language Models (LLMs) AI Frontpage 2025 Top Fifty: 14 % 154 Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers by Sam Marks , Adam Karvonen , James Chua , Subhash Kantamneni , Euan Ong , Julian Minder , Clément Dumas , Owain_Evans 18th Dec 2025 AI Alignment Forum Linkpost for arxiv.org 10 min read 11 154 Ω 73 TL;DR: We train LLMs to accept LLM neural activations as inputs and answer arbitrary questions about them in natural l

Explore this link on the map →

related reading