flâneur — a map of the web's best reading

Vanessa Kosoy's Shortform — LessWrong

lesswrong.com · 570 words · saved by 1 readers

Comment by Vanessa Kosoy - Epistemic status: half-baked Arguably, an aligned AI should be aligned to the user's prior as well as the user's utility function. Hence, any value-learning protocol should also be doing prior-learning. The problem is, any learning process requires (explicitly or implicitly) its own prior. But shouldn't this also be the user's prior? Is this an infinite regress? Maybe not: here is a way out that seems elegant in a way. For now, we will work in the Bayesian framework. Let be the set of possible universes (the domain of our beliefs). Let be the kernel s.t. is the prior of the user that lives in universe (i.e. the prior of the user, assuming universe is the real universe). Let be the likelihood function of the evidence observed by the AI so far. Then, the AI might wish to find a prior s.t. updating on evidence and then taking the posterior-expectation of the user's prior yields again. That is, we propose for the AI to choose a prior which is in some sense self-endorsing. Mathematically, this is saying that there should be constant s.t. So, needs to be an eigenvector of the operator on the left hand side! Imagining for simplicity that is a finite set, assuming (reasonably) that is always non-dogmatic (i.e. for all ) and that is non-zero everywhere (otherwise the universes where it is zero can be discarded first), the Perron-Frobenius theorem implies that a unique positive eigenvector exists. This seems compelling! The problem is, this doesn't describe a Bayesian agent: as the AI accumulates more evidence, its prior changes and hence its belief changes in a non-Bayesian way. Maaaybe this is some kind of "radical probabilism" (I don't understand the latter well enough to say). From a different angle, what I really want is a priorist ("updateless") specification of the agent's policy, and atm I don't know how to reconcile it with this "eigenprior". Also, this feels to good to be true: we get a canonical prior out of nothing? This brings to mind the sort of negative

Explore this link on the map →

saved by