flâneur

I Built TetrisBench, Where LLMs Compete at Playing Tetris. Here’s What I Found. | Andreessen Horowitz

a16z.com · 1,729 words · saved by 1 readers

Turning Tetris into a coding and optimization loop shows how GPT-5.2, Claude Opus 4.5, Gemini 3, Grok 4.1, and Sonnet 4 differ in long-horizon reasoning, strategy adaptation, intervention timing, and behavior under edge cases and shifting state.

I was playing Tetris99 one evening on my Nintendo Switch and started wondering whether I could build a version where I play against an LLM. Not in the sense of proving that a model could “beat” the game, but simply to see what it would feel like to play something familiar against a system that reasons very differently from a human. At the time, I wasn’t trying to create a benchmark. I didn’t have a leaderboard to climb, or a claim to make. I was mostly curious about the experience itself since Tetris is deceptively simple. At its core, it’s an optimization problem: every move is a tradeoff…

saved by

related reading