flâneur

Notes on Inference Integrity - by James Tillman - ForeWord

newsletter.forethought.org · 763 words · saved by 1 readers

Claude Fable’s deliberately triggered sandbagging shows that training-time targets are, by themselves, insufficient to guarantee particular LLM behaviors.

This article was created by Forethought. See all our research on our website. Summary: Claude Fable’s deliberately triggered sandbagging shows that training-time targets are, by themselves, insufficient to guarantee particular LLM behaviors. To preserve the public’s reasonable confidence in LLM behaviors, LLM foundation model companies should take inference-time guarantees as seriously as their model specs. When the system card for Anthropic’s Fable was published on June 9th, the card noted that using Fable for “frontier LLM development” would run contrary to the terms of service for the…

saved by

related reading