How enabling two settings tripled our scores on the ARC-AGI-3 benchmark | OpenAI
A sped-up video of GPT‑5.6 Sol attempting to solve puzzles in the ARC-AGI-3 benchmark, with the official harness (left) and our Responses API harness (right), which retains reasoning and enables compaction. On the leaderboard for this game (opens in a new window) , no frontier model solves any level beyond the first. With our harness, GPT‑5.6 Sol solves all six. When we first saw GPT‑5.6 Sol’s low scores on the ARC-AGI-3 (opens in a new window) benchmark, we were puzzled. GPT‑5.6 Sol has solved longstanding open problems in mathematics like the cycle double cover conjecture (opens in a new window) and beaten games like Pokémon FireRed. But on ARC-AGI-3, a benchmark of 2D puzzle games, GPT‑5.6 Sol scored just 7.8%, and GPT‑5.5 could barely play the games at all, scoring a paltry 0.4%. Were 2D puzzle games unusually difficult for our models? Or was something else going on? Benchmarks rarely measure AI models in isolation. They also measure less visible choices about API settings, ha