Consequentialist cognition — LessWrong
Consequentialist reasoning selects policies on the basis of their predicted consequences - it does action X because X is forecasted to lead to preferred outcome Y. Whenever we reason that an agent which prefers outcome Y over Y′ will therefore do X instead of X′, we're implicitly assuming that the agent has the cognitive ability to do consequentialism at least about Xs and Ys. It does means-end reasoning; it selects means on the basis of their predicted ends plus a preference over ends. E.g: When we infer that a paperclip maximizer would try to improve its own cognitive abilities given means to do so, the background assumptions include: * That the paperclip maximizer can forecast the consequences of the policies "self-improve" and "don't try to self-improve"; * That the forecasted consequences are respectively "more paperclips eventually" and "less paperclips eventually"; * That the paperclip maximizer preference-orders outcomes on the basis of how many paperclips they contain; * That the paperclip maximizer outputs the immediate action it predicts will lead to more future paperclips. (Technically, since the forecasts of our actions' consequences will usually be uncertain, a coherent agent needs a utility function over outcomes and not just a preference ordering over outcomes.) The related idea of "backward chaining" is one particular way of solving the cognitive problems of consequentialism: start from a desired outcome/event/future, and figure out what intermediate events are likely to have the consequence of bringing about that event/outcome, and repeat this question until it arrives back at a particular plan/policy/action. Many narrow AI algorithms are consequentialists over narrow domains. A chess program that searches far ahead in the game tree is a consequentialist; it outputs chess moves based on the expected result of those chess moves and your replies to them, into the distant future of the board. We can see one of the critical aspects of human i
x Consequentialist cognition — LessWrong Consequentialist cognition Edited by Eliezer Yudkowsky , et al. last updated 11th Jun 2016 Consequentialist reasoning selects policies on the basis of their predicted consequences - it does action X because X is forecasted to lead to preferred outcome Y . Whenever we reason that an agent which prefers outcome Y over Y ′ will therefore do X instead of X ′ , we're implicitly assuming that the agent has the cognitive ability to do consequentialism at least about X s and Y s. It does means-end reasoning; it selects means on the basis of their predicted ends
Explore this link on the map →related reading
- Why Tool AIs Want to Be Agent AIs · Gwern.netgwern.net
- Instrumental convergence — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- Beliefs are Chosen to Serve Goals — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Varieties Of Doomminihf.com
- ⿻ Symbiogenesis vs. Convergent Consequentialism — LessWronglesswrong.com
- Why AIs aren't power-seeking yet — LessWronglesswrong.com
- The Best of LessWrong — LessWronglesswrong.com
- [1902.09469] Embedded Agencyarxiv.org
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com