Metrics — Continual Learning Bench Docs
Each task is designed to answer one question: does the system perform better the more experience it has? Each task defines a per-instance reward. Tasks have their own definition (accuracy, normalized error, profit, …), but every reward must satisfy: If per-instance difficulty is roughly homogeneous across the sequence, reward by itself is enough to read off learning: a system that is genuinely improving from experience will produce an upward-trending reward curve, because the curve isn't being moved by changes in instance difficulty. In that regime, "higher reward at task instance t" means "more was learned by instance t." For example, in the "exploitable poker" task, chips are reset after each hand (so winnings are not accumulated). That means that when a system is playing against the same opponent, the basic challenge for that hand is the same as in previous hands. Increases in reward across hands can thus reflect real learning signal — systems are doing progressively better against
Metrics — Continual Learning Bench Docs
Explore this link on the map →related reading
- Concepts — Continual Learning Bench Docscontinual-learning-bench.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- The Era of Experience Paper.pdfstorage.googleapis.com
- EdgeBench | Scaling Laws of Environment Learningedge-bench.org
- Learning Beyond Gradientstrinkle23897.github.io
- Why We Need Continual Learning | Andreessen Horowitza16z.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- [2605.12474] Reward Hacking in Rubric-Based Reinforcement Learningarxiv.org
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- Learning curve - Wikipediaen.wikipedia.org
- DataRater: Meta-Learned Dataset Curationarxiv.org