How to reason from first principles – Casey Handmer's blog
This post is a follow up on my general questions on AI post, and expands on ideas I published here and here in the ancient times. Every time one of the labs releases an updated model I give it a thorough shakedown on physics, in the style of the oral examination that is still used in Europe and a few other places. Claude, Grok, Gemini, and GPT are all advancing by leaps and bounds on a wide variety of evals, some of which include rather advanced or technical questions in both math and science, including Physics Olympiad-style problems, or grad school qualifying exams. And yet, none of these models would be able to pass the physicist Turing test. It’s not even a matter of knowledge, I know of reasonably talented middle schoolers with no specialized physics training who could reason and infer on some of these basic questions in a much more fluent and intuitive way. For context, I competed in the IPhO in 2005, and still have a copy of my notes (they are short, rough, and not 100% accurate
This post is a follow up on my general questions on AI post , and expands on ideas I published here and here in the ancient times. Every time one of the labs releases an updated model I give it a thorough shakedown on physics, in the style of the oral examination that is still used in Europe and a few other places. Claude, Grok, Gemini, and GPT are all advancing by leaps and bounds on a wide variety of evals, some of which include rather advanced or technical questions in both math and science, including Physics Olympiad-style problems , or grad school qualifying exams . And yet, none of these
saved by
related reading
- As Rocks May Think | Eric Jangevjang.com
- Machina Mirabilis - Michael Hlamichaelhla.com
- Vibe physics: The AI grad student \ Anthropicanthropic.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- How To Understand Things - Nabeel S. Qureshinabeelqu.substack.com
- principia labsprincipialabs.org
- Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazinequantamagazine.org
- Can activation verbalizers surface an internal chain of thought? — LessWronglesswrong.com
- Physics of Language Modelsphysics.allen-zhu.com
- Mathematics in the Library of Babel - Daniel Littdaniellitt.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- GenAI Handbookgenai-handbook.github.io