Teaching ML to answer questions honestly instead of predicting human answers | by Paul Christiano | AI Alignment
In this post I consider the problem of models learning “predict how a human would answer questions” instead of “answer questions honestly.” (A special case of the problem from Inaccessible Information.) I describe a possible three-step approach for learning to answer questions honestly instead: I don’t know whether this problem is a relatively unimportant special case of alignment, or one of the core difficulties. In any case, my next step will be trying to generate failure stories that definitely cannot be addressed by any of the angles of attack I know so far (including the ones in this post). I think it’s relatively unlikely that almost anything specific I said here will really hold up over the long term, but I do think I’ve learned something about each of these steps. If the ideas end up being important then you can expect a future post with a simpler algorithm, more confidence that it works, clearer definitions, and working code. (Thanks to Ajeya Cotra, David Krueger, and Mark Xu
Teaching ML to answer questions honestly instead of predicting human answers Paul Christiano 20 min read · May 28, 2021 -- Listen Share In this post I consider the problem of models learning “predict how a human would answer questions” instead of “answer questions honestly.” (A special case of the problem from Inaccessible Information .) I describe a possible three-step approach for learning to answer questions honestly instead: Change the learning process so that it does not have a strong inductive bias towards “predict human answers,” by allowing the complexity of the honest question-answeri
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- How confessions can keep language models honest | OpenAIopenai.com
- Teaching Claude why \ Anthropicanthropic.com
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Unsupervised Elicitationalignment.anthropic.com
- Mediumai-alignment.com
- Models don’t seem to be dishonest in the way humans are — LessWronglesswrong.com
- Alignment faking in large language modelsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com