flâneur — a map of the web's best reading

News — BenchCAD

benchcad.com · 431 words · saved by 1 readers

Execution-grounded benchmark for LLMs & multimodal models on programmatic CAD (CadQuery).

2026-07-09 OpenAI evaluated its GPT-5.6 family on BenchCAD OpenAI's GPT-5.6 launch post lists BenchCAD in its computer-use table, next to OSWorld and BrowseComp — the second frontier lab to report on the benchmark, and the first to run it across a whole model family. As published: GPT-5.6 Sol scores 70.6% without tools and 83.4% with a Python tool , with Terra at 62.3% / 78.2%, Luna at 63.1% / 73.9%, and GPT-5.5 at 44.4% / 55.8%. Scores exactly as printed in OpenAI's launch table (computer-use section, 2026-07-09). None are re-graded by BenchCAD. The Claude bars are Anthropic's own published v

Explore this link on the map →

saved by

related reading