flâneur — a map of the web's best reading

Dimitris Papailiopoulos on X: "ECHO: Terminal Agents Learn World Models for Free" / Twitter

x.com · saved by 1 readers

To view keyboard shortcuts, press question mark View keyboard shortcuts Home Notifications Chat Bookmarks Articles Profile More Tweet Tim Kostolansky @thkostolansky Tweet See new posts Conversation Dimitris Papailiopoulos @DimitrisPapail ECHO: Terminal Agents Learn World Models for Free 45 179 844 399K Co-written with @VaishShrivas We taught CLI agents to predict terminal responses during RL, alongside the usual GRPO loss on actions. The change is tiny: same rollout and forward pass, but stop masking out terminal-output tokens. The effect is huge: all evals improve, and the resulting models measurably learn how the terminal behaves. CLI agents can learn a terminal model for free — and use it to act better! This is ECHO: a hybrid objective that trains on both sides of the interaction: what the agent writes, and what the terminal writes back. Check out the full paper, and code on top of SkyRL. GIF If you’re too busy to read this whole post, here’s what we found: Standard agent RL throws

Explore this link on the map →

saved by