How well do truth probes generalise? — LessWrong
Representation engineering (RepEng) has emerged as a promising research avenue for model interpretability and control. Recent papers have proposed me…
x How well do truth probes generalise? — LessWrong Activation Engineering AI Frontpage 96 How well do truth probes generalise? by mishajw 24th Feb 2024 11 min read 11 96 Representation engineering (RepEng) has emerged as a promising research avenue for model interpretability and control. Recent papers have proposed methods for discovering truth in models with unlabeled data , guiding generation by modifying representations , and building LLM lie detectors . RepEng asks the question: If we treat representations as the central unit, how much power do we have over a model’s behaviour? Most techni
saved by
related reading
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- 2310.01405arxiv.org
- Actually, Othello-GPT Has A Linear Emergent World Representation - Neel Nandaneelnanda.io
- Structure and Interpretation of Deep Networkssidn.baulab.info
- [2310.06824] The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasetsarxiv.org
- 2212.03827arxiv.org
- truthfulQA_lin_evans.pdfowainevans.github.io
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- [2603.07267] How to Steal Reasoning Without Reasoning Tracesarxiv.org
- 3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitationalignment.anthropic.com