LLMs and World Models, Part 1 - by Melanie Mitchell
In the long-ago times, before large-scale generative AI came on the scene, machine-learning systems had some problems: often they didn’t learn the general concepts we were trying to teach them, but rather solved problems using “shortcuts” or “surface heuristics.” To give one stark example, some researchers tried to train a deep neural network to classify skin lesions in photos like the one below as “benign” or “malignant.” While this network performed well on the kinds of photos it was trained on, the researchers noticed a problem: “[T]he algorithm appeared more likely to interpret images with rulers as malignant. Why? In our dataset, images with rulers were more likely to be malignant; thus, the algorithm inadvertently ‘learned’ that rulers are malignant.” In short, the network learned a useful heuristic: if there are features corresponding to a ruler in the image, the answer is “malignant.” The network didn’t understand what a lesion is, or what motivated the researchers to train it,
LLMs and World Models, Part 1 How do Large Language Models Make Sense of Their “Worlds”? Melanie Mitchell Feb 13, 2025 313 17 24 Share This is part 1 of a two-part post on LLMs and “world models.” Part 2 is here . AI Brittleness in the Before Times In the long-ago times, before large-scale generative AI came on the scene, machine-learning systems had some problems: often they didn’t learn the general concepts we were trying to teach them, but rather solved problems using “shortcuts” or “surface heuristics.” To give one stark example , some researchers tried to train a deep neural network to cl
Explore this link on the map →saved by
related reading
- Language Models, World Models, and Human Model-Buildinglingo.csail.mit.edu
- World Models: Computing the Uncomputablenotboring.co
- Large Language Model: world models or surface statistics?thegradient.pub
- GenAI Handbookgenai-handbook.github.io
- LLM Daydreaming · Gwern.netgwern.net
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- pdfopenreview.net
- Explore | alphaXivalphaxiv.org
- Against LLM Reductionism — LessWronglesswrong.com
- True Agents Model the Worldprimeintellect.ai
- Language Models in Plato's Cave - by Sergey Levinesergeylevine.substack.com
- World Models | Rohit Bandarurohitbandaru.github.io