flâneur — a map of the web's best reading

Is This Lie Detector Really Just a Lie Detector? An Investigation of LLM Probe Specificity. — LessWrong

lesswrong.com · 6,086 words · saved by 1 readers

Whereas previous work has focused primarily on demonstrating a putative lie detector’s sensitivity/generalizability[1][2], it is equally important to evaluate its specificity. With this in mind, I evaluated a lie detector trained with a state-of-the-art, white box technique - probing an LLM’s activations during production of facts/lies - and found that it had high sensitivity but low specificity. The detector might be better thought of as identifying when the LLM is doing something other than fact-based retrieval (e.g. when writing fiction), which spans a much wider surface area than it should. I found that the detector could be made more specific through data augmentation, but that this improved specificity did not transfer to other domains, unfortunately. I hope that this study sheds light on some of the remaining gaps in our tooling for and understanding of lie detection - and probing more generally - and points in directions toward improving them. You can find the associated cod

x Is This Lie Detector Really Just a Lie Detector? An Investigation of LLM Probe Specificity. — LessWrong Eliciting Latent Knowledge Interpretability (ML & AI) Language Models (LLMs) AI Frontpage 43 Is This Lie Detector Really Just a Lie Detector? An Investigation of LLM Probe Specificity. by Josh Levy 4th Jun 2024 AI Alignment Forum 22 min read 0 43 Ω 20 Abstract Whereas previous work has focused primarily on demonstrating a putative lie detector’s sensitivity/generalizability [1] [2] , it is equally important to evaluate its specificity. With this in mind, I evaluated a lie detector trained

Explore this link on the map →

related reading