3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitation
We study 3 realistic challenges to the safety of unsupervised elicitation and easy-to-hard generalization techniques, which aim to steer models on tasks which are beyond human supervision. We create datasets to test the robustness of methods against these challenges. We stress-test existing techniques on them along with new methods relying on 2 hopes: ensembling and combining unsupervised and easy-to-hard methods. We find that although the new hopes sometimes perform better than other approaches, no technique reliably performs well on the 3 challenges. 📝Paper, 💻Code Research done as part of MATS and the Anthropic fellowship. To steer language models towards truthful outputs on tasks which are beyond human capability, previous work has suggested training models on easy tasks to steer them on harder ones (easy-to-hard generalization), or using unsupervised training algorithms to steer models with no external labels at all (unsupervised elicitation). In this new work, we Most easy-to-ha
3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitation Alignment Science Blog 3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitation Callum Canavan 1,2 , Aditya Shrivastava 1,2 March 5, 2026 Allison Qi 3 , Jonathan Michala 1 Fabien Roger 3 1 MATS; 2 Anthropic Fellows Program; 3 Anthropic tl;dr We study 3 realistic challenges to the safety of unsupervised elicitation and easy-to-hard generalization techniques, which aim to steer models on tasks which are beyond human supervision. We create datasets to test the robustness of methods against these challenges. We stress-t
saved by
related reading
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- Unsupervised Elicitationalignment.anthropic.com
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Unsupervised Elicitation of Language Modelsarxiv.org
- 2212.03827arxiv.org
- How well do truth probes generalise? — LessWronglesswrong.com
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Foundation Models for Oversight | Transluce AItransluce.org
- Large Language Models Must Be Taught to Know What They Don't Knowarxiv.org
- Language Models Learn to Mislead Humans via RLHFarxiv.org
- [2602.05910] Chunky Post-Training: Data Driven Failures of Generalizationarxiv.org