Blueprint-Bench: Testing spatial intelligence in AI models | Andon Labs
How do AI models understand space? We test this by asking them to convert apartment photographs into accurate 2D floor plans. While photos are familiar training data, spatial reconstruction requires genuine intelligence.
Eval How do AI models understand space? We test this by asking them to convert apartment photographs into accurate 2D floor plans. While photos are familiar training data, spatial reconstruction requires genuine intelligence. Leaderboard Model Type Similarity Score (mean) 1 Human* Human 0.547 2 GPT-5 LLM 0.431 3 Gemini 2.5 Pro LLM 0.421 4 GPT-5-mini LLM 0.400 5 Grok-4 LLM 0.393 6 Codex CLI (GPT-5) Agent 0.388 7 Gemini 2.5 Flash LLM 0.362 8 Claude Code (Opus 4) Agent 0.355 9 GPT Image Image Model 0.300 10 Claude Opus 4 LLM 0.286 11 Random baseline…
saved by
related reading
- LLM Visualizationbbycroft.net
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Replicate - Run AI with an APIreplicate.com
- The Shape of AI | UX Patterns for Artificial Intelligence Designshapeof.ai
- There's An AI For That® — The front page of AItheresanaiforthat.com
- Together AI | The AI Native Cloudtogether.ai
- AI Agent Benchmark for Real-World Professional Workflowsagents-last-exam.org
- MAI-Thinking-1 | Microsoft AImicrosoft.ai
- Explore | alphaXivalphaxiv.org
- justinzwu.comjustinzwu.com
- Cosmoscosmos.so
- Contra Labs - Powered by Contracontralabs.com