flâneur

Scaling Agentic RL: 365,000+ Environments for SWE, Terminal, and Search

primeintellect.ai · 2,694 words · saved by 1 readers

We ship first-party integrations for 23 agentic tasksets behind one taskset API - ~365,000 software engineering, terminal, and search tasks, validated and ready for evals and RL training on Prime Intellect infrastructure.

Scaling Agentic RL: 365,000+ Environments for SWE, Terminal, and Search The open research ecosystem has produced many great datasets for the three main agentic domains - software engineering, terminal use, and web research - but every one of them ships with its own harness, its own image conventions, its own grading scripts, and its own failure modes. 23 tasksets behind one taskset API DOMAIN · TASKSETSTASKS SOFTWARE ENGINEERING swesmith 83,519openswe 36,884swerebench_v2 32,079scaleswe 17,202swelego 15,903multiswe 6,835r2e_gym 4,578swebench_pro 731swebench_verified…

saved by

related reading