I Built TetrisBench, Where LLMs Compete at Playing Tetris. Here’s What I Found. | Andreessen Horowitz
Turning Tetris into a coding and optimization loop shows how GPT-5.2, Claude Opus 4.5, Gemini 3, Grok 4.1, and Sonnet 4 differ in long-horizon reasoning, strategy adaptation, intervention timing, and behavior under edge cases and shifting state.
I was playing Tetris99 one evening on my Nintendo Switch and started wondering whether I could build a version where I play against an LLM. Not in the sense of proving that a model could “beat” the game, but simply to see what it would feel like to play something familiar against a system that reasons very differently from a human. At the time, I wasn’t trying to create a benchmark. I didn’t have a leaderboard to climb, or a claim to make. I was mostly curious about the experience itself since Tetris is deceptively simple. At its core, it’s an optimization problem: every move is a tradeoff…
saved by
related reading
- As Rocks May Think | Eric Jangevjang.com
- LLM Chess: Benchmarking Reasoning and Instruction-Following in LLMs through Chessarxiv.org
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- combinatorial reasoning environments for LLMs and RLdjdumpling.github.io
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Composer2.pdfcursor.com
- Learning to reason with LLMs | OpenAIopenai.com
- Playing chess with large language modelsnicholas.carlini.com
- Explore | alphaXivalphaxiv.org
- @jeremy-berman/arc-agi on Paramsparams.com
- Can you be sure to clear a line at Tetris? - a3nm's bloga3nm.net
- 2025: The year in LLMssimonwillison.net