WebArena-Infinity: Generating Browser Environments with Verifiable Tasks at Scale
Realistic web environments with verifiable tasks are essential for training and evaluating general-purpose browser-use agents. However, constructing such environments remains prohibitively labor-intensive, requiring manual effort across sandbox creation, data curation, task design, and programmatic verifier authoring1234. The required expertise, often including significant programming effort, further limits scalability, confining large-scale environment construction largely to well-resourced organizations. We aim to automate what has traditionally been a manual, expert-driven process. In this work, we propose an approach for automatically generating high-authenticity, high-complexity, and RL-ready environments, along with verifiable tasks, from static artifacts such as recorded workflows and user manuals using a multi-agent system. The resulting pipeline is both time- and cost-efficient: each environment can be generated within ten hours at a cost under $100, and the process is highly
GitHub Environment Hub Dataset browser-use agents navigating generated environments across diverse application domains. Realistic web environments with verifiable tasks are essential for training and evaluating general-purpose browser-use agents. However, constructing such environments remains prohibitively labor-intensive, requiring manual effort across sandbox creation, data curation, task design, and programmatic verifier authoring1234. The required expertise, often including significant programming effort, further limits scalability, confining large-scale environment construction…
saved by
related reading
- Akash Bajwa on X: "RL Environments with Scale AI" / Xx.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- Mechanize, Inc.mechanize.work
- General Agent: A Self-Evolving, Synthetic Agent Environmentprimeintellect.ai
- Scaling Agentic RL: 365,000+ Environments for SWE, Terminal, and Searchprimeintellect.ai
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- General Agent: A Self-Evolving, Synthetic Agent Environmentprimeintellect.ai
- Enabling Agent 3 to Self-Test at Scale with REPL-Based Verificationblog.replit.com
- Harness design for long-running application developmentanthropic.com
- RL environment creation is becoming continuous QA - Rishi Desairishidesai.org
- EdgeBench | Scaling Laws of Environment Learningedge-bench.org