
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark | OpenAI
OpenAI discovered that enabling retained reasoning and compaction settings tripled GPT-5.6 Sol's scores on the ARC-AGI-3 benchmark. This indicates that harness design and API settings significantly influence model performance results.
Why it matters
Developers can improve AI efficiency and accuracy by using production settings rather than generic harnesses. It warns that some benchmarks may underrepresent a model's actual capabilities.
The details
- GPT-5.6 Sol's score increased from 13.3% to 38.3% on the public set.
- Output tokens were reduced by 6x using retained reasoning and compaction.
- The official harness discarded private reasoning and used rolling truncation of history.
Show entities and relationshipsHide entities and relationships
In this article
Technologies
Companies
Organizations
Key connections
ARC Prize developed and released the ARC-AGI-3 benchmark
GPT-5.6 Sol uses Responses API
GPT-5.6 Sol was evaluated using Responses API with retained reasoning and compaction
ChatGPT uses Responses API
ChatGPT uses Responses API features including retained reasoning and compaction
Codex uses Responses API
Codex uses Responses API features including retained reasoning and compaction
GPT-5.6 Sol is related to ARC-AGI-3
GPT-5.6 Sol was benchmarked on ARC-AGI-3
GPT-5.5 is related to ARC-AGI-3
GPT-5.5 was tested on ARC-AGI-3
Show 4 more connectionsShow fewer connections
Responses API is built with Context Compaction
Responses API includes context compaction capability
Responses API is built with Retained Reasoning
Responses API supports retained reasoning across turns
OpenAI owns Responses API
OpenAI provides the Responses API developer platform
GPT-5.6 Sol competes with GPT-5.5
GPT-5.6 Sol outperformed GPT-5.5 while operating at lower reasoning effort
Related events
GPT-5.6 Sol Scores Tripled on ARC-AGI-3 Benchmark
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.