flâneur — a map of the web's best reading

I Built TetrisBench, Where LLMs Compete at Playing Tetris. Here’s What I Found. | Andreessen Horowitz

a16z.com · saved by 1 readers

Turning Tetris into a coding and optimization loop shows how GPT-5.2, Claude Opus 4.5, Gemini 3, Grok 4.1, and Sonnet 4 differ in long-horizon reasoning, strategy adaptation, intervention timing, and behavior under edge cases and shifting state.

Explore this link on the map →

saved by