flâneur — a map of the web's best reading

Inverse Rubric Optimization: A testbed for agent science | Fulcrum

fulcrum.inc · 5,212 words · saved by 2 readers

We propose inverse rubric optimization (IRO): tasks where an agent must learn the preferences of a black-box judge under a label budget. IRO tasks induce rich agent behavior and smooth scaling, making them a useful testbed for agent science.

“It is important to draw wisdom from many different places. If you take it from only one place, it becomes rigid and stale.” — Uncle Iroh At Fulcrum Research, we study the performance and behavior of long-horizon agents. Although each task setting has its own specific structure, we believe it’s possible to find general principles of agent performance across settings, each contributing to a nascent agent science. In this post, we motivate the difficulty of finding suitable settings for agent science and propose inverse rubric optimization (IRO) settings, in which an agent has to optimize the pr

Explore this link on the map →

saved by

related reading