✳flâneur — a map of the web's best reading
EdgeBench | Scaling Laws of Environment Learning
edge-bench.org · 2,575 words · saved by 1 readers
EdgeBench studies how agents learn from real-world environments across 134 day-long executable tasks.
EdgeBench studies how agents learn from real-world environments across 134 day-long executable tasks. Most benchmarks score what a model already knows. EdgeBench is built to measure something else. It asks how an agent when it is given the time, the feedback, and the room to improve. Every workspace, feedback signal, and judge approximates real practice, so a high score reflects what an agent Each task runs 12+ hours of continuous operation, long enough for experience to compound. Selected extended runs continue Tasks span science, software engineering, optimization, knowledge work, formal mat
Explore this link on the map →saved by
related reading
- The Era of Experience Paper.pdfstorage.googleapis.com
- Composer2.pdfcursor.com
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era | Qiuyang Mangjoyemang33.github.io
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Machine Studying | Jacob Xiaochen Lijacobxli.com
- Learning Beyond Gradientstrinkle23897.github.io
- PostTrainBenchposttrainbench.com
- Can AI Learn From Experience? EBR-Bench Results | Epoch AI | Epoch AIepoch.ai