✳flâneur — a map of the web's best reading
ceselder/cot-oracle: Our improvements to activation oracles ·
github.com · 1,140 words · saved by 1 readers
Our improvements to activation oracles
CoT Oracle CoT Oracle is a white-box chain-of-thought monitor built on Activation Oracles. The core model is Qwen3-8B with a LoRA adapter trained to read its own residual-stream activations and answer questions about the reasoning that produced them. This README documents the pipeline the current code actually runs. Older docs elsewhere in the repo, especially under src/evals/ , describe older or parallel experiments and are not the main training-time path. Core Mechanism A source sequence is built from a question plus a chain-of-thought or other context text. Activation positions are chosen f
Explore this link on the map →related reading
- cot-oracle/AObench at main · ceselder/cot-oracle · GitHubgithub.com
- Current activation oracles are hard to use — LessWronglesswrong.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- [2606.02609] Building Better Activation Oraclesarxiv.org
- [2606.02609] Building Better Activation Oraclesarxiv.org
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWronglesswrong.com
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersalignment.anthropic.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- [2512.15674] Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersarxiv.org
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Composer2.pdfcursor.com