[2302.02083] Theory of Mind May Have Spontaneously Emerged in Large Language Models
Theory of mind (ToM), or the ability to impute unobservable mental states to others, is central to human social interactions, communication, empathy, self-consciousness, and morality. We tested several language models using 40 classic false-belief tasks widely used to test ToM in humans. The models published before 2020 showed virtually no ability to solve ToM tasks. Yet, the first version of GPT-3 ("davinci-001"), published in May 2020, solved about 40% of false-belief tasks-performance comparable with 3.5-year-old children. Its second version ("davinci-002"; January 2022) solved 70% of false-belief tasks, performance comparable with six-year-olds. Its most recent version, GPT-3.5 ("davinci-003"; November 2022), solved 90% of false-belief tasks, at the level of seven-year-olds. GPT-4 published in March 2023 solved nearly all the tasks (95%). These findings suggest that ToM-like ability (thus far considered to be uniquely human) may have spontaneously emerged as a byproduct of language models' improving language skills.
[2302.02083] Evaluating Large Language Models in Theory of Mind Tasks Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2302.02083 (cs) [Submitted on 4 Feb 2023 ( v1 ), last revised 4 Nov 2024 (this version, v7)] Title: Evaluating Large Language Models in Theory of Mind Tasks Authors: Michal Kosinski View a PDF of the paper titled Evaluating Large Language Models in Theory of Mind Tasks, by Michal Kosinski View PDF Abstract: Eleven Large Language Models
Explore this link on the map →related reading
- [2302.02083] Theory of Mind May Have Spontaneously Emerged in Large Language Modelsarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- AI hype is built on high test scores. Those tests are flawed. | MIT Technology Reviewtechnologyreview.com
- Do large language models understand us? | by Blaise Aguera y Arcas | Mediummedium.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- [2406.02061] Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Modelsarxiv.org
- Large Language Model: world models or surface statistics?thegradient.pub
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Cognitive Biases in Large Language Models — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Theory of Mind for Multi-Agent Collaboration via Large Language Modelsarxiv.org
- Language Models can Solve Computer Tasksarxiv.org