EdgeBench | Scaling Laws of Environment Learning
edge-bench.org · 2,575 words · saved by 2 readers
EdgeBench studies how agents learn from real-world environments across 134 day-long executable tasks.
EdgeBench studies how agents learn from real-world environments across 134 day-long executable tasks. Most benchmarks score what a model already knows. EdgeBench is built to measure something else. It asks how an agent when it is given the time, the feedback, and the room to improve. Every workspace, feedback signal, and judge approximates real practice, so a high score reflects what an agent Each task runs 12+ hours of continuous operation, long enough for experience to compound. Selected extended runs continue Tasks span science, software engineering, optimization, knowledge work, formal mat
saved by
related reading
- ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering | alphaXivalphaxiv.org
- The Era of Experience Paper.pdfstorage.googleapis.com
- Can AI Learn From Experience? EBR-Bench Results | Epoch AI | Epoch AIepoch.ai
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- FrontierSWEfrontierswe.com
- Scaling Agentic RL: 365,000+ Environments for SWE, Terminal, and Searchprimeintellect.ai
- Mercor to acquire Deeptunemercor.com
- PostTrainBenchposttrainbench.com
- Machine Studying | Jacob Xiaochen Lijacobxli.com
- Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era | Qiuyang Mangjoyemang33.github.io
- Datacurve | The data engine for frontier AIdatacurve.ai
- FrontierSWEfrontierswe.com