flâneur

Manan Khattar

0 followers · 2 following · 233 views

on the atlas — 4

highlights — 3

  • Telescope uses short, evidence-grounded facets inspired by the same bottom-up approach as Clio [7]. For traces, it reconstructs the session from user messages, visible agent replies, tool calls, errors, metadata, and other context. From there, Telescope summarizes what went wrong, embeds those summaries, and uses density-based clustering [8] to form emergent issue groups.
    Michele Catasta on X: "Continual Learning for Agents" / X
  • The system has two measurement pillars and one optimization loop. Offline benchmarks tell us whether candidate changes can complete simulated app-building tasks before we ship them. Online A/B tests and production traces show how real users are affected after the changes ship. Those signals then flow back into evals and shipping decisions.
    Michele Catasta on X: "Continual Learning for Agents" / X
  • This is where the a massive (but often overlooked) opportunity lies: harness-level learning lets you mine production traces to systematically improve the code, tools, and instructions that power every instance of your agent, while context-level learning lets you personalize at the agent, user, and org level, so your product gets better with every interaction. Do all the above, and you will be compounding improvements that you can ship daily.
    Michele Catasta on X: "Continual Learning for Agents" / X