How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
Ranking
Overall
68
Content
75
Popularity
N/A
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - OpenAI reports that two API settings—reasoning retention and compaction—tripled GPT-5.6’s ARC-AGI-3 benchmark score while improving efficiency.
- Retaining reasoning context helped the model preserve useful intermediate work.
- Compaction reduced context overhead during longer benchmark interactions.
- The result highlights how inference configuration can materially affect benchmark performance.
- The provided summary does not specify exact scores or API parameter names.
Sources (1)
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
Public signals
N/A
TL;DR - OpenAI reports that two API settings—reasoning retention and compaction—tripled GPT-5.6’s ARC-AGI-3 benchmark score while improving efficiency.
- Retaining reasoning context helped the model preserve useful intermediate work.
- Compaction reduced context overhead during longer benchmark interactions.
- The result highlights how inference configuration can materially affect benchmark performance.
- The provided summary does not specify exact scores or API parameter names.