flâneur — a map of the web's best reading

CausaLab — Can LLM Agents Discover Causal Mechanisms by Experiment?

dylanzsz.github.io · 2,610 words · saved by 1 readers

A scalable environment that scores not just whether an agent gets the answer, but whether it recovered the causal mechanism — by experimenting like a scientist.

CausaLab — Can LLM Agents Discover Causal Mechanisms by Experiment? TL;DR CausaLab evaluates two things at once : did the agent solve the task, and is its answer grounded in a faithful recovered causal mechanism ? Each episode hides a freshly sampled structural causal model (SCM), so you can't win by reciting memorized causal facts. Prediction ≠ understanding. On observational 6-node graphs, GPT-5.2-high reaches 92% task accuracy but only 0.47 all-edge F 1 . The right number, the wrong graph. How you experiment is the whole game. Observation narrows the hypothesis space; agent-chosen intervent

Explore this link on the map →

saved by

related reading