MazeBench Results — Maze Bench Blog
mazebench.com · 2,978 words · saved by 1 readers
MazeBench benchmark results.
Introducing MazeBench MazeBench is an enormous open world resembling an intense labyrinth, constructed to address the capabilities of long running agent loops inside three-dimensional space. It evaluates whether agents demonstrate enough visual spatial reasoning to perceive their surroundings and execute elaborate plans. To make MazeBench more interesting, Sokoban-style box pushing puzzles were placed throughout the world. The boxes come in a variety of shapes and sizes, so agents must correctly interpret their structure before they can plan how to use them. These puzzles require reasoning…
saved by
related reading
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- As Rocks May Think | Eric Jangevjang.com
- FrontierSWEfrontierswe.com
- Explore | alphaXivalphaxiv.org
- Kimi K2.5 Tech Blog: Visual Agentic Intelligencekimi.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- TERMINAL-BENCHtbench.ai
- Datacurve | The data engine for frontier AIdatacurve.ai
- EdgeBench | Scaling Laws of Environment Learningedge-bench.org
- FrontierSWEfrontierswe.com
- PostTrainBenchposttrainbench.com
- combinatorial reasoning environments for LLMs and RLdjdumpling.github.io