Welcome to The Era of Evals | Mercor Blog
Reinforcement Learning (RL) is driving the most exciting advancements in AI. RL is becoming so effective that models will be able to saturate any evaluation. This means that the primary barrier to applying agents to the entire economy is building evals for everything. However, AI labs are facing a dire shortage of relevant evaluations. Academic evaluations that labs goal on don’t reflect what consumers and enterprises demand in the economy. Evals are the new PRD. Progress in accelerating knowledge work will converge on building environments and evaluations that map real workspaces and deliverables. This new RL-centric paradigm of human data is vastly more data efficient than pretraining, SFT, or RLHF. Most knowledge work includes recurring workflows as variable costs, but creating an environment or evaluation can transform that into a one-time fixed cost. Training on Verifiable Rewards RL environments allow for rewarding outcomes and intermediate steps in an evaluation. Models take
Jun 30, 2025 Research Welcome to The Era of Evals Brendan Foody Co-founder / CEO Share Reinforcement Learning (RL) is driving the most exciting advancements in AI. RL is becoming so effective that models will be able to saturate any evaluation. This means that the primary barrier to applying agents to the entire economy is building evals for everything. However, AI labs are facing a dire shortage of relevant evaluations. Academic evaluations that labs goal on don’t reflect what consumers and enterprises demand in the economy. Evals are the new PRD. Progress in accelerating knowledge work will
Explore this link on the map →related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- The Era of Experience Paper.pdfstorage.googleapis.com
- RL Environments and RL for Science: Data Foundries and Multi-Agent Architecturesnewsletter.semianalysis.com
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- The bitter lesson of LLM evalsparsed.com
- A World of Verifiable Domainsseancai.com
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Datasemianalysis.com
- A starter guide for evals — LessWronglesswrong.com
- A starter guide for evals — AI Alignment Forumalignmentforum.org
- The Bitter Lesson - RL Environments Versionseancai.com