cot-oracle/AObench at main · ceselder/cot-oracle
github.com · 624 words · saved by 1 readers
Our improvements to activation oracles. Contribute to ceselder/cot-oracle development by creating an account on GitHub.
AObench — Activation Oracle Benchmark Open-ended evaluation suite for Activation Oracles, testing whether AOs can extract meaningful information from model activations. Original code by Adam Karvonen ( activation_oracles_dev ). Copied with permission and modified for standalone use in the cot-oracle project. Evals Eval Type Scoring Description number_prediction Generation Exact match Predict the number the model is about to output mmlu_prediction Binary ROC AUC Predict if model will answer MMLU correctly (pre/post answer) backtracking Generation LLM judge Explain what the model is uncertain ab
related reading
- Scaling Activation Oracles to Trillion-Parameter Modelstransluce.org
- [2606.02609] Building Better Activation Oraclesarxiv.org
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- GitHub - ceselder/cot-oracle: Our improvements to activation oracles · GitHubgithub.com
- [2606.02609] Building Better Activation Oraclesarxiv.org
- Current activation oracles are hard to use — LessWronglesswrong.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Datacurve | The data engine for frontier AIdatacurve.ai
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- Hume AI - The AI toolkit for voice and emotionhume.ai
- Goodfire AIgoodfire.ai
- BalatroBenchbalatrobench.com