Disagreeable Me: A Multi-Level view of LLM Intentionality
Prompted by Keith Frankish's recent streamed discussion of LLM intentionality on YouTube, there's a particular idea I wanted to share which I'm not sure is widely enough appreciated but which I think gives a valuable perspective from which to think about LLMs and what kinds of intentions they may have. This is not an original idea of my own -- at least some of my thinking on this was sparked by reading about the Waluigi problem on LessWrong. In this post, I'm going to be making the case that LLMs might have intentionality much like ours, but please understand that I'm making a point in principle and not so much arguing for the capabilities of current LLMs, which are probably not there yet. I'm going to be talking about what scope there is to give the benefit of the doubt to arbitrarily competent future LLMs, albeit ones that follow more or less the same paradigms as those of today. I'm going to try to undermine some proposed reasons for skepticism about the intentions or understanding
Disagreeable Me: A Multi-Level view of LLM Intentionality Tuesday, 9 May 2023 A Multi-Level view of LLM Intentionality Prompted by Keith Frankish's recent streamed discussion of LLM intentionality on YouTube , there's a particular idea I wanted to share which I'm not sure is widely enough appreciated but which I think gives a valuable perspective from which to think about LLMs and what kinds of intentions they may have. This is not an original idea of my own -- at least some of my thinking on this was sparked by reading about the Waluigi problem on LessWrong. In this post, I'm going to be maki
saved by
related reading
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Prompt Injection as Role Confusionrole-confusion.github.io
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWronglesswrong.com
- LLMs Get Lost in Evolving User Intentarxiv.org
- Against LLM Reductionism — LessWronglesswrong.com
- Don't dethrone consciousness! - by Erik Hoeltheintrinsicperspective.com
- No, Artificial Intelligence Is Not Conscious - The Atlantictheatlantic.com
- No, Artificial Intelligence Is Not Conscious - The Atlantictheatlantic.com
- the void — LessWronglesswrong.com