Scaling Agentic RL: 365,000+ Environments for SWE, Terminal, and Search
We ship first-party integrations for 23 agentic tasksets behind one taskset API - ~365,000 software engineering, terminal, and search tasks, validated and ready for evals and RL training on Prime Intellect infrastructure.
Scaling Agentic RL: 365,000+ Environments for SWE, Terminal, and Search The open research ecosystem has produced many great datasets for the three main agentic domains - software engineering, terminal use, and web research - but every one of them ships with its own harness, its own image conventions, its own grading scripts, and its own failure modes. 23 tasksets behind one taskset API DOMAIN · TASKSETSTASKS SOFTWARE ENGINEERING swesmith 83,519openswe 36,884swerebench_v2 32,079scaleswe 17,202swelego 15,903multiswe 6,835r2e_gym 4,578swebench_pro 731swebench_verified…
saved by
related reading
- General Agent: A Self-Evolving, Synthetic Agent Environmentprimeintellect.ai
- How Zapier Turned AutomationBench Into a Continuous Agent Improvement Loopprimeintellect.ai
- General Agent: A Self-Evolving, Synthetic Agent Environmentprimeintellect.ai
- SWE-1.7: Frontier Intelligence at a Fraction of the Costcognition.com
- Senior SWE-Benchsenior-swe-bench.snorkel.ai
- Introducing SWE-grep and SWE-grep-mini: RL for Multi-Turn, Fast Context Retrieval | Cognitioncognition.ai
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Composer2.pdfcursor.com
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- FrontierSWEfrontierswe.com
- RL Environments and RL for Science: Data Foundries and Multi-Agent Architecturesnewsletter.semianalysis.com
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work