flâneur — a map of the web's best reading

EdgeBench | Scaling Laws of Environment Learning

edge-bench.org · 2,575 words · saved by 1 readers

EdgeBench studies how agents learn from real-world environments across 134 day-long executable tasks.

EdgeBench studies how agents learn from real-world environments across 134 day-long executable tasks. Most benchmarks score what a model already knows. EdgeBench is built to measure something else. It asks how an agent when it is given the time, the feedback, and the room to improve. Every workspace, feedback signal, and judge approximates real practice, so a high score reflects what an agent Each task runs 12+ hours of continuous operation, long enough for experience to compound. Selected extended runs continue Tasks span science, software engineering, optimization, knowledge work, formal mat

Explore this link on the map →

saved by

related reading