✳flâneur — a map of the web's best reading
cot-oracle/AObench at main · ceselder/cot-oracle
github.com · 624 words · saved by 1 readers
Our improvements to activation oracles. Contribute to ceselder/cot-oracle development by creating an account on GitHub.
AObench — Activation Oracle Benchmark Open-ended evaluation suite for Activation Oracles, testing whether AOs can extract meaningful information from model activations. Original code by Adam Karvonen ( activation_oracles_dev ). Copied with permission and modified for standalone use in the cot-oracle project. Evals Eval Type Scoring Description number_prediction Generation Exact match Predict the number the model is about to output mmlu_prediction Binary ROC AUC Predict if model will answer MMLU correctly (pre/post answer) backtracking Generation LLM judge Explain what the model is uncertain ab
Explore this link on the map →related reading
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- GitHub - ceselder/cot-oracle: Our improvements to activation oracles · GitHubgithub.com
- Current activation oracles are hard to use — LessWronglesswrong.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- [2606.02609] Building Better Activation Oraclesarxiv.org
- Cookbookcookbook.openai.com
- [2606.02609] Building Better Activation Oraclesarxiv.org
- GitHub - SoyGema/pulling_ace · GitHubgithub.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- AI Benchmark Leaderboards & Model Evals | BenchmarkListbenchmarklist.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com