3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitation
We study 3 realistic challenges to the safety of unsupervised elicitation and easy-to-hard generalization techniques, which aim to steer models on tasks which are beyond human supervision. We create datasets to test the robustness of methods against these challenges. We stress-test existing techniques on them along with new methods relying on 2 hopes: ensembling and combining unsupervised and easy-to-hard methods. We find that although the new hopes sometimes perform better than other approaches, no technique reliably performs well on the 3 challenges. 📝Paper, 💻Code Research done as part of MATS and the Anthropic fellowship. To steer language models towards truthful outputs on tasks which are beyond human capability, previous work has suggested training models on easy tasks to steer them on harder ones (easy-to-hard generalization), or using unsupervised training algorithms to steer models with no external labels at all (unsupervised elicitation). In this new work, we Most easy-to-ha
3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitation Alignment Science Blog 3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitation Callum Canavan 1,2 , Aditya Shrivastava 1,2 March 5, 2026 Allison Qi 3 , Jonathan Michala 1 Fabien Roger 3 1 MATS; 2 Anthropic Fellows Program; 3 Anthropic tl;dr We study 3 realistic challenges to the safety of unsupervised elicitation and easy-to-hard generalization techniques, which aim to steer models on tasks which are beyond human supervision. We create datasets to test the robustness of methods against these challenges. We stress-t
Explore this link on the map →saved by
related reading
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- Unsupervised Elicitationalignment.anthropic.com
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- How well do truth probes generalise? — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- Eliciting Language Model Behaviors with Investigator Agents | Transluce AItransluce.org
- The Unreasonable Effectiveness of Easy Training Data for Hard Tasksalexandrabarr.beehiiv.com
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com