OpenAI demonstrated that enabling retained reasoning and compaction in the Responses API tripled GPT-5.6 Sol's score on ARC-AGI-3 from 13.3% to 38.3% while reducing output tokens by 6x.
Jul 29, 2026
2d agoKey Details
- GPT-5.6 Sol scored 13.3% with the official generic harness
- Score increased to 38.3% on the public task set when using Responses API harness with retained reasoning and compaction
- Output token usage was reduced by 6x
- GPT-5.5 scored 0.4% on the benchmark