Standards for Belief Representations in LLMs
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. As large language models (LLMs) continue to demonstrate remarkable abilities across various domains, computer scientists are developing methods to understand their cognitive processes, particularly concerning how (and if) LLMs internally represent their beliefs about the world. However, this field currently lacks a unified theoretical foundation to underpin the study of belief in LLMs. This article begins filling this gap by proposing adequacy conditions for a representation in an LLM to count as belief-like. We argue that, while the project of belief measurement in LLMs shares striking features with belief measurement as carried out in decision theory and formal epistemology, it also differs in ways that should change how we measure belief. Thus, drawin
Standards for Belief Representations in LLMs Daniel A. Herrmann, Benjamin A. Levinstein Abstract. As large language models (LLMs) continue to demonstrate remarkable abilities across various domains, computer scientists are developing methods to understand their cognitive processes, particularly concerning how (and if) LLMs internally represent their beliefs about the world. However, this field currently lacks a unified theoretical foundation to underpin the study of belief in LLMs. This article begins filling this gap by proposing adequacy conditions for a representation in an LLM to count as
Explore this link on the map →related reading
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?alignment.anthropic.com
- Large Language Model: world models or surface statistics?thegradient.pub
- Against LLM Reductionism — LessWronglesswrong.com
- [2310.06824] The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasetsarxiv.org
- How well do truth probes generalise? — LessWronglesswrong.com
- The bitter lesson of LLM evalsparsed.com
- The Future of Everything is Lies, I Guessaphyr.com
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- Teaching LLMs to reason like Bayesiansresearch.google
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Do large language models understand us? | by Blaise Aguera y Arcas | Mediummedium.com