flâneur

How Zapier Turned AutomationBench Into a Continuous Agent Improvement Loop

primeintellect.ai · 549 words · saved by 1 readers

Zapier built AutomationBench on Prime Intellect Lab to evaluate agents on realistic multi-step automation workflows, catch reward-hacking behavior in live metrics, and turn evals into RL training environments.

At a glance Built AutomationBench for real-world, multi-step automation workflows. Tested agents across Zapier's 9,000+ app ecosystem. Prime Intellect Lab caught reward hacking live: API fetch calls dropped to near zero while reward stayed flat. Debugged eval quality through rollout traces and training metrics. Turned evals into RL training environments. Moved from static benchmarks to continuous agent improvement. Ran RL experiments without managing infra manually. Background Zapier connects over 9,000 apps into automated workflows, helping teams replace repetitive manual work with…

saved by

related reading