How well do truth probes generalise? — LessWrong
Representation engineering (RepEng) has emerged as a promising research avenue for model interpretability and control. Recent papers have proposed me…
x How well do truth probes generalise? — LessWrong Activation Engineering AI Frontpage 96 How well do truth probes generalise? by mishajw 24th Feb 2024 11 min read 11 96 Representation engineering (RepEng) has emerged as a promising research avenue for model interpretability and control. Recent papers have proposed methods for discovering truth in models with unlabeled data , guiding generation by modifying representations , and building LLM lie detectors . RepEng asks the question: If we treat representations as the central unit, how much power do we have over a model’s behaviour? Most techni
Explore this link on the map →saved by
related reading
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- Actually, Othello-GPT Has A Linear Emergent World Representation - Neel Nandaneelnanda.io
- [2310.06824] The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasetsarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Trustworthy AI: Validity, Fairness, Explainability, and Uncertainty Assessments: Explainability methods: Linear Probescarpentries-incubator.github.io
- Is This Lie Detector Really Just a Lie Detector? An Investigation of LLM Probe Specificity. — LessWronglesswrong.com
- 3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitationalignment.anthropic.com
- Detecting Strategic Deception Using Linear Probes — LessWronglesswrong.com
- Coup probes: Catching catastrophes with probes trained off-policy — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com