flâneur — a map of the web's best reading

Teaching ML to answer questions honestly instead of predicting human answers | by Paul Christiano | AI Alignment

ai-alignment.com · 4,987 words · saved by 1 readers

In this post I consider the problem of models learning “predict how a human would answer questions” instead of “answer questions honestly.” (A special case of the problem from Inaccessible Information.) I describe a possible three-step approach for learning to answer questions honestly instead: I don’t know whether this problem is a relatively unimportant special case of alignment, or one of the core difficulties. In any case, my next step will be trying to generate failure stories that definitely cannot be addressed by any of the angles of attack I know so far (including the ones in this post). I think it’s relatively unlikely that almost anything specific I said here will really hold up over the long term, but I do think I’ve learned something about each of these steps. If the ideas end up being important then you can expect a future post with a simpler algorithm, more confidence that it works, clearer definitions, and working code. (Thanks to Ajeya Cotra, David Krueger, and Mark Xu

Teaching ML to answer questions honestly instead of predicting human answers Paul Christiano 20 min read · May 28, 2021 -- Listen Share In this post I consider the problem of models learning “predict how a human would answer questions” instead of “answer questions honestly.” (A special case of the problem from Inaccessible Information .) I describe a possible three-step approach for learning to answer questions honestly instead: Change the learning process so that it does not have a strong inductive bias towards “predict human answers,” by allowing the complexity of the honest question-answeri

Explore this link on the map →

related reading