Can risk aversion learned at low stakes generalize to astronomically high stakes?
This post covers our recent paper: Out-of-Distribution Generalization of Risk Aversion in Language Models. It gives the intro, main results, and example prompts from the training and evaluation sets. For everything else, see the paper.
This post covers our recent paper: Out-of-Distribution Generalization of Risk Aversion in Language Models. It gives the intro, main results, and example prompts from the training and evaluation sets. For everything else, see the paper. Training AIs to be risk-averse in resources could be a useful failsafe in case of misalignment. Misaligned but risk-averse AIs would tend to prefer a higher chance of modest payments to a lower chance of successful rebellion, so in many circumstances we could pay these AIs to cooperate with us. But we can only feasibly train AIs to be risk-averse on…
saved by
related reading
- Risk-Averse AIsforethought.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Teaching Claude why \ Anthropicanthropic.com
- Many arguments for AI x-risk are wrong — AI Alignment Forumalignmentforum.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- How far does alignment midtraining generalize?alignment.openai.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Foundation Models for Oversight | Transluce AItransluce.org
- AI Safety | Arkosevictoriabrook.github.io