flâneur — a map of the web's best reading

How well do truth probes generalise? — LessWrong

lesswrong.com · 5,116 words · saved by 2 readers

Representation engineering (RepEng) has emerged as a promising research avenue for model interpretability and control. Recent papers have proposed me…

x How well do truth probes generalise? — LessWrong Activation Engineering AI Frontpage 96 How well do truth probes generalise? by mishajw 24th Feb 2024 11 min read 11 96 Representation engineering (RepEng) has emerged as a promising research avenue for model interpretability and control. Recent papers have proposed methods for discovering truth in models with unlabeled data , guiding generation by modifying representations , and building LLM lie detectors . RepEng asks the question: If we treat representations as the central unit, how much power do we have over a model’s behaviour? Most techni

Explore this link on the map →

saved by

related reading