flâneur — a map of the web's best reading

Introducing Terminal-Bench 2.0 and Harbor

tbench.ai · 355 words · saved by 1 readers

A benchmark for terminal agents

Today we are releasing Terminal-Bench 2.0 and Harbor: a harder, better verified version of Terminal-Bench and a new package for evaluating and optimizing agents. Harbor While building Terminal-Bench we kept hearing about the same set of problems from agent developers. Namely: Evaluating in containers is slow, how can we scale horizontally to thousands of containers in the cloud? How can we not only evaluate but also improve agents via SFT, RL, and prompt optimization? With so many frameworks for building agents and benchmarks for measuring them, how do we build tools that generalize across dep

Explore this link on the map →

related reading