How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
TL;DR - OpenAI reports that two API settings—reasoning retention and compaction—tripled GPT-5.6’s ARC-AGI-3 benchmark score while improving efficiency.
- Retaining reasoning context helped the model preserve useful intermediate work.
- Compaction reduced context overhead during longer benchmark interactions.
- The result highlights how inference configuration can materially affect benchmark performance.
- The provided summary does not specify exact scores or API parameter names.