Audit Without Verification: When LLM Accountability Layers Relay Rather Than Check
Multi-agent LLM pipelines increasingly span organisational boundaries, and when a fault surfaces someone must determine where it entered. In deployment the artifact available for that determination is rarely a full execution trace: it is the reports each agent filed, and a filed report can state a conclusion alongside its observations. We ask what an accountability layer built on such reports can and cannot do. Using a pre-registered, institutionally partitioned pipeline of six agents with process-level information boundaries, exactly balanced defect injection and matched clean twins (345,600 requests per chain model, two models), we first report that our pre-registered hypothesis — that collective responsibility framing degrades escalation increasingly with chain length — is not supported. The layer nevertheless fails, and it fails asymmetrically. It originates almost nothing: zero allegations across 7,996 clean episodes where every agent stayed silent. It filters upstream error poorl
Abstract Multi-agent LLM pipelines increasingly span organisational boundaries, and when a fault surfaces someone must determine where it entered. In deployment the artifact available for that determination is rarely a full execution trace: it is the reports each agent filed, and a filed report can state a conclusion alongside its observations. We ask what an accountability layer built on such reports can and cannot do. Using a pre-registered, institutionally partitioned pipeline of six agents with process-level information boundaries, exactly balanced defect injection and matched clean…
saved by
related reading
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Building Effective AI Agents \ Anthropicanthropic.com
- Building Effective AI Agents \ Anthropicanthropic.com
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWronglesswrong.com
- Agent Observability and Tracingarize.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com
- Prompt Injection as Role Confusionrole-confusion.github.io
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationsimonwillison.net
- LLM-as-a-Verifier: A General-Purpose Verification Framework | alphaXivalphaxiv.org