Is This Lie Detector Really Just a Lie Detector? An Investigation of LLM Probe Specificity. — LessWrong
Whereas previous work has focused primarily on demonstrating a putative lie detector’s sensitivity/generalizability[1][2], it is equally important to evaluate its specificity. With this in mind, I evaluated a lie detector trained with a state-of-the-art, white box technique - probing an LLM’s activations during production of facts/lies - and found that it had high sensitivity but low specificity. The detector might be better thought of as identifying when the LLM is doing something other than fact-based retrieval (e.g. when writing fiction), which spans a much wider surface area than it should. I found that the detector could be made more specific through data augmentation, but that this improved specificity did not transfer to other domains, unfortunately. I hope that this study sheds light on some of the remaining gaps in our tooling for and understanding of lie detection - and probing more generally - and points in directions toward improving them. You can find the associated cod
x Is This Lie Detector Really Just a Lie Detector? An Investigation of LLM Probe Specificity. — LessWrong Eliciting Latent Knowledge Interpretability (ML & AI) Language Models (LLMs) AI Frontpage 43 Is This Lie Detector Really Just a Lie Detector? An Investigation of LLM Probe Specificity. by Josh Levy 4th Jun 2024 AI Alignment Forum 22 min read 0 43 Ω 20 Abstract Whereas previous work has focused primarily on demonstrating a putative lie detector’s sensitivity/generalizability [1] [2] , it is equally important to evaluate its specificity. With this in mind, I evaluated a lie detector trained
Explore this link on the map →related reading
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- How well do truth probes generalise? — LessWronglesswrong.com
- Detecting Strategic Deception Using Linear Probes — LessWronglesswrong.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- The Road To Honest AI - by Scott Alexanderastralcodexten.com
- [2310.06824] The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasetsarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Simple probes can catch sleeper agents \ Anthropicanthropic.com