flâneur — a map of the web's best reading

LLMs can learn about themselves by introspection — LessWrong

lesswrong.com · 11,345 words · saved by 1 readers

TLDR: We find that LLMs are capable of introspection on simple tasks. We discuss potential implications of introspection for interpretability and the moral status of AIs. Paper Authors: Felix Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, Owain Evans This post contains edited extracts from the full paper. Humans acquire knowledge by observing the external world, but also by introspection. Introspection gives a person privileged access to their current state of mind (e.g., thoughts and feelings) that is not accessible to external observers. Can LLMs introspect? We define introspection as acquiring knowledge that is not contained in or derived from training data but instead originates from internal states. Such a capability could enhance model interpretability. Instead of painstakingly analyzing a model's internal workings, we could simply ask the model about its beliefs, world models, and goals. More speculatively, an introspect

x LLMs can learn about themselves by introspection — LessWrong AI Frontpage 110 LLMs can learn about themselves by introspection by Felix J Binder , Owain_Evans 18th Oct 2024 AI Alignment Forum 11 min read 38 110 Ω 40 Are LLMs capable of introspection, i.e. special access to their own inner states? Can they use this access to report facts about themselves that are not in the training data? Yes — in simple tasks at least! TLDR: We find that LLMs are capable of introspection on simple tasks. We discuss potential implications of introspection for interpretability and the moral status of AIs. Pape

Explore this link on the map →

related reading