flâneur

SHIFT relies on token-level features to de-bias Bias in Bios probes — LessWrong

lesswrong.com · saved by 1 readers

In Sparse Feature Circuits (Marks et al. 2024), the authors introduced Spurious Human-Interpretable Feature Trimming (SHIFT), a technique designed to eliminate unwanted features from a model's computational process. They validate SHIFT on the Bias in Bios task, which we think is too simple to serve as meaningful validation. To summarize: We don’t think the results in this post show that SHIFT is a bad method, but rather that the Bias in Bios dataset (or any other simple dataset) is not a good test bed to judge SHIFT or other similar methods. Follow-ups of SHIFT-like methods (e.g., Karvonen et al. (2025), Casademunt et al. (2025), SAE Bench) have already used more complex datasets and found promising results. However, these studies still focus on fairly toy settings, and we are not aware of research that focuses on disentangling safety-relevant concepts (e.g., sycophancy versus correctness in reward models)[1]. In the rest of the post, we give some background on the method, share our ex

saved by