Notes on Inference Integrity - by James Tillman - ForeWord
Claude Fable’s deliberately triggered sandbagging shows that training-time targets are, by themselves, insufficient to guarantee particular LLM behaviors.
This article was created by Forethought. See all our research on our website. Summary: Claude Fable’s deliberately triggered sandbagging shows that training-time targets are, by themselves, insufficient to guarantee particular LLM behaviors. To preserve the public’s reasonable confidence in LLM behaviors, LLM foundation model companies should take inference-time guarantees as seriously as their model specs. When the system card for Anthropic’s Fable was published on June 9th, the card noted that using Fable for “frontier LLM development” would run contrary to the terms of service for the…
saved by
related reading
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- confessions_paper.pdfcdn.openai.com
- Thoughts on Claude Fable's silent safeguards — LessWronglesswrong.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com