Agents Can Get Stuck in Self-distrusting Equilibria — LessWrong
This post was written as part of research done at MATS 9.0 under the mentorship of Richard Ngo. He contributed significantly to the ideas discussed within. This post questions the sanctity of the "agent" and discusses how Temporal Instances (TIs) of an agent can enter conflict due to distrust. These dynamics are describable mathematically as an intrapersonal cooperative game. I define a time-version of Nash equilibria and show an example of a self-punishing pattern between TIs that is nevertheless stable. This leads us to ask what conditions allow disparate parts of an agent to cooperate harmoniously. I conjecture that agents showing a degree of consistency in their actions over time can be seen as adhering to an identity that replaces Common Knowledge of Rationality (CKR) between the game's players. In subscribing to a common identity, TIs declare trust in each other akin to that which an updateless[1] agent would embody. I next deliberate on the shape that a formal statement and proo