flâneur — a map of the web's best reading

3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitation

alignment.anthropic.com · 2,971 words · saved by 1 readers

We study 3 realistic challenges to the safety of unsupervised elicitation and easy-to-hard generalization techniques, which aim to steer models on tasks which are beyond human supervision. We create datasets to test the robustness of methods against these challenges. We stress-test existing techniques on them along with new methods relying on 2 hopes: ensembling and combining unsupervised and easy-to-hard methods. We find that although the new hopes sometimes perform better than other approaches, no technique reliably performs well on the 3 challenges. 📝Paper, 💻Code Research done as part of MATS and the Anthropic fellowship. To steer language models towards truthful outputs on tasks which are beyond human capability, previous work has suggested training models on easy tasks to steer them on harder ones (easy-to-hard generalization), or using unsupervised training algorithms to steer models with no external labels at all (unsupervised elicitation). In this new work, we Most easy-to-ha

3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitation Alignment Science Blog 3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitation Callum Canavan 1,2 , Aditya Shrivastava 1,2 March 5, 2026 Allison Qi 3 , Jonathan Michala 1 Fabien Roger 3 1 MATS; 2 Anthropic Fellows Program; 3 Anthropic tl;dr We study 3 realistic challenges to the safety of unsupervised elicitation and easy-to-hard generalization techniques, which aim to steer models on tasks which are beyond human supervision. We create datasets to test the robustness of methods against these challenges. We stress-t

Explore this link on the map →

saved by

related reading