OpenAI has published a post examining why GPT-5.6 Sol scored so poorly on ARC-AGI-3, an interactive reasoning benchmark[1]. The cause was not the model's capability but the settings of the harness that runs the evaluation. Turning on two API settings the company already uses in ChatGPT and Codex roughly tripled the score on the public task set and cut output tokens by a factor of 6.

A benchmark that drops agents into unfamiliar games

ARC-AGI-3 is an interactive reasoning benchmark built by ARC Prize. Agents are placed inside 2D games they have never seen, with no instructions, and have to infer the rules by acting, identify a goal on their own, and pursue it[1][2]. Rather than solving static puzzles, the benchmark measures whether a model can learn from experience over time, and the 25 public demo games are all playable by humans[1][2].

On this benchmark GPT-5.6 Sol scored just 7.8 percent. The previous generation, GPT-5.5, could barely play the games at all and landed at 0.4 percent[1]. The verified scores published by ARC Prize show the gap clearly: Sol's maximum reasoning setting reaches 96.5 percent on ARC-AGI-1 and 92.5 percent on ARC-AGI-2, but only 7.78 percent on ARC-AGI-3[3].

The numbers looked odd from OpenAI's side as well. The company notes that GPT-5.6 Sol has worked on longstanding open problems such as the cycle double cover conjecture, beat Pokémon FireRed with a vision-only harness, and cleared Slay the Spire using Codex computer use[1]. It is hard to believe that 2D puzzle games alone would be an outsized weakness.

Reasoning was discarded every turn, and history was truncated

When OpenAI inspected the ARC-AGI-3 harness, it found two behaviors[1].

First, all of the model's private reasoning was discarded after every game action. A record of past moves and short notes remained, but the plans and insights behind those moves were gone, so the model had to work the game out from scratch each turn. Second, the harness used a rolling truncation window that dropped the oldest messages once the context exceeded 175,000 characters. That meant the model lost not only its earlier thinking but eventually its earlier actions too[1].

The generic design is deliberate. ARC's reasoning is that a simple harness with no tools or special features makes model shortcomings more visible and comparisons fairer. Commercial developers, by contrast, optimize the harness around each model's features and quirks[1]. That difference showed up directly in the score.

Rebuilding on the Responses API let learning accumulate

OpenAI's models are trained to write private reasoning messages before producing replies or tool calls, and to keep those messages as part of the conversation history, summarizing the history when it grows too long. To match that production setup, the company reimplemented the ARC-AGI-3 harness on its Responses API. With GPT-5.6, passing the previous response ID automatically retains reasoning across tool calls and turns[1].

Two changes followed. The model spent less time thinking before each action, because it no longer had to reinterpret the game every turn, and with access to its earlier thoughts it learned over time and stuck to coherent strategies[1].

The next step was replacing rolling truncation with compaction. Compaction is a Responses API feature that, instead of simply dropping old history, carries forward the state and reasoning that matter in a condensed form so fewer tokens are needed[4]. Truncation has two drawbacks: earlier observations and actions are lost, and the model spends much of the run operating with a nearly full context window, which can slightly impair performance. With compaction enabled, Sol preserved what it had learned about each game across longer runs and reached a higher score with fewer output tokens[1].

Combined, retained reasoning and compaction gave GPT-5.6 Sol at its maximum reasoning setting roughly 3 times the score with 6 times fewer output tokens, according to OpenAI[1].

Evaluation numbers do not measure the model alone

The argument boils down to one point: benchmarks rarely measure models in isolation. They also measure the less visible choices around API settings, harness design, and prompting. OpenAI adds that this is not the first time it has been surprised by a low public benchmark score only to find that the eval runner used a generic harness that dropped reasoning messages[1].

Its recommendation to API developers is straightforward: use the Responses API rather than the legacy Chat Completions API, retain reasoning, and enable compaction. When comparing models, the company suggests relying on evals that use those settings, since they best match how the models are actually deployed in ChatGPT and Codex[1].

OpenAI notes that the investigation was inspired by ARC's own analysis of GPT-5.5's shortcomings, and thanks ARC for its years of work on AGI evaluation[1]. Given that ARC's harness is an intentional design choice aimed at fairness, the takeaway is less about who is right and more about checking what is held fixed and what is variable when reading a score.

Summary

The main reason GPT-5.6 Sol scored low on ARC-AGI-3 was a harness that discarded private reasoning after every action and truncated older history. Enabling two API settings, retained reasoning and compaction, roughly tripled the score and cut output tokens by a factor of 6. Benchmark figures do not reflect raw model ability alone, and they need to be read together with which API and which settings produced them.

Source[1]: https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/

Source[2]: https://arcprize.org/arc-agi/3

Source[3]: https://arcprize.org/results/openai-gpt-5-6-sol

Source[4]: https://developers.openai.com/api/docs/guides/compaction