flâneur

Can risk aversion learned at low stakes generalize to astronomically high stakes?

substack.com · 1,290 words · saved by 1 readers

This post covers our recent paper: Out-of-Distribution Generalization of Risk Aversion in Language Models. It gives the intro, main results, and example prompts from the training and evaluation sets. For everything else, see the paper.

This post covers our recent paper: Out-of-Distribution Generalization of Risk Aversion in Language Models. It gives the intro, main results, and example prompts from the training and evaluation sets. For everything else, see the paper. Training AIs to be risk-averse in resources could be a useful failsafe in case of misalignment. Misaligned but risk-averse AIs would tend to prefer a higher chance of modest payments to a lower chance of successful rebellion, so in many circumstances we could pay these AIs to cooperate with us. But we can only feasibly train AIs to be risk-averse on…

saved by

related reading