Your Agents Are Not Time Aware — LessWrong
We had two CLI agents, Claude Code and Codex, predict, execute and then retrospectively estimate their own wall-clock runtime. We ran experiments across ProgramBench, our own evaluation-suite (AgentTime), and then ablated on what information agents have access to. As agents become more and more capable, we turn to measure the length of tasks they can automate. METR’s time-horizon curves show that frontier models are able to solve 17 hour tasks 50% of the time. But this only measures time externally. For an agent to be able to work for certain very long-horizon tasks, they would need a calibrated model of its temporality. An agent needs to know how long it’s been working, how much time is left for a budget or deadline, and make policy decisions based on it. An agent can spend time without any representation of time passing. We ourselves are also bad at containing good representations of time. For example, human are often susceptible to biases such as the planning fallacy, the holiday pa