Physics of Language Models
Even today, GPT-4 and Llama-3 still provide incorrect answers to some questions that are simple for humans. Is this a problem inherent to GPT-4, or is it due to insufficient training? Is its mathematical capability too weak? Does this only affect models as of July 2024, or will GPT-6 and Llama-5 also face this problem? What about other pre-trained models? (Spoiler alert: these are counterexamples that all today's LLMs shall fail — a.k.a. Turing tests.) Pretrained LLMs (like GPT-4, LLaMA-3, Claude-3) are like monkeys used in animal behavior science, and we now live in an era where most people can interact with these monkeys and play games. This is fantastic! However, rigorous scientists need to think about the underlying "why" to uncover the "universal laws" behind these phenomena, rather than merely studying individual monkeys. People often celebrate when an LLM ranks high on a benchmark, but is this truly accurate? Could the models have "seen" this benchmark data during training? Imag
Physics of Language Models Search this site Embedded Files Skip to main content Skip to navigation Physics of Language Models The concept of Physics of Language Models was jointly conceived and designed by ZA and Xiaoli Xu. Watch Order (or Not) All videos are self-contained . You may watch them in any order , or selectively, depending on your interests . What follows is a suggested navigation, not a prerequisite path. Tutorial I = Part 1 (structure) + Part 2 (reasoning) + Part 3 (knowledge) ← ← ICML 2024 Tutorial (2 hours). A high-level synthesis of results from earlier Parts 1+2+3 of the seri
Explore this link on the map →related reading
- Physics of Language Modelsphysics.allen-zhu.com
- gpt-4.pdfcdn.openai.com
- GPT-4openai.com
- Language Models, World Models, and Human Model-Buildinglingo.csail.mit.edu
- Large Language Model: world models or surface statistics?thegradient.pub
- Things we learned about LLMs in 2024simonwillison.net
- larger language models may disappoint you [or, an eternally unfinished draft] — LessWronglesswrong.com
- Against LLM Reductionism — LessWronglesswrong.com
- Catching up on the weird world of LLMssimonwillison.net
- What will GPT-2030 look like? — AI Alignment Forumalignmentforum.org
- 2310.10631arxiv.org
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com