✳flâneur — a map of the web's best reading
Introducing Terminal-Bench 2.0 and Harbor
tbench.ai · 355 words · saved by 1 readers
A benchmark for terminal agents
Today we are releasing Terminal-Bench 2.0 and Harbor: a harder, better verified version of Terminal-Bench and a new package for evaluating and optimizing agents. Harbor While building Terminal-Bench we kept hearing about the same set of problems from agent developers. Namely: Evaluating in containers is slow, how can we scale horizontally to thousands of containers in the cloud? How can we not only evaluate but also improve agents via SFT, RL, and prompt optimization? With so many frameworks for building agents and benchmarks for measuring them, how do we build tools that generalize across dep
Explore this link on the map →related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- Composer2.pdfcursor.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- benchmarks.bio — Agentic AI benchmarks on messy, real-world biological databenchmarks.bio
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- AI Agent Benchmark for Real-World Professional Workflowsagents-last-exam.org
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- GitHub - microsoft/WindowsAgentArena: Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents. · GitHubgithub.com
- Through the looking glass of benchmark hacking — Poolsidepoolside.ai
- Agent Observability and Tracingarize.com