flâneur — a map of the web's best reading

Physics of Language Models

physics.allen-zhu.com · 1,152 words · saved by 1 readers

Even today, GPT-4 and Llama-3 still provide incorrect answers to some questions that are simple for humans. Is this a problem inherent to GPT-4, or is it due to insufficient training? Is its mathematical capability too weak? Does this only affect models as of July 2024, or will GPT-6 and Llama-5 also face this problem? What about other pre-trained models? (Spoiler alert: these are counterexamples that all today's LLMs shall fail — a.k.a. Turing tests.) Pretrained LLMs (like GPT-4, LLaMA-3, Claude-3) are like monkeys used in animal behavior science, and we now live in an era where most people can interact with these monkeys and play games. This is fantastic! However, rigorous scientists need to think about the underlying "why" to uncover the "universal laws" behind these phenomena, rather than merely studying individual monkeys. People often celebrate when an LLM ranks high on a benchmark, but is this truly accurate? Could the models have "seen" this benchmark data during training? Imag

Physics of Language Models Search this site Embedded Files Skip to main content Skip to navigation Physics of Language Models The concept of Physics of Language Models was jointly conceived and designed by ZA and Xiaoli Xu. Watch Order (or Not) All videos are self-contained . You may watch them in any order , or selectively, depending on your interests . What follows is a suggested navigation, not a prerequisite path. Tutorial I = Part 1 (structure) + Part 2 (reasoning) + Part 3 (knowledge) ← ← ICML 2024 Tutorial (2 hours). A high-level synthesis of results from earlier Parts 1+2+3 of the seri

Explore this link on the map →

related reading