Inverse Rubric Optimization: A testbed for agent science | Fulcrum
We propose inverse rubric optimization (IRO): tasks where an agent must learn the preferences of a black-box judge under a label budget. IRO tasks induce rich agent behavior and smooth scaling, making them a useful testbed for agent science.
“It is important to draw wisdom from many different places. If you take it from only one place, it becomes rigid and stale.” — Uncle Iroh At Fulcrum Research, we study the performance and behavior of long-horizon agents. Although each task setting has its own specific structure, we believe it’s possible to find general principles of agent performance across settings, each contributing to a nascent agent science. In this post, we motivate the difficulty of finding suitable settings for agent science and propose inverse rubric optimization (IRO) settings, in which an agent has to optimize the pr
Explore this link on the map →saved by
related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Composer2.pdfcursor.com
- Building Effective AI Agents \ Anthropicanthropic.com
- PostTrainBenchposttrainbench.com
- Agent Observability and Tracingarize.com
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- AI Agent Benchmark for Real-World Professional Workflowsagents-last-exam.org
- [2605.12474] Reward Hacking in Rubric-Based Reinforcement Learningarxiv.org
- Machine Studying | Jacob Xiaochen Lijacobxli.com
- The Verification Horizon: No Silver Bullet for Coding Agent Rewardsarxiv.org
- GitHub - METR/RE-Bench · GitHubgithub.com