How two API settings tripled GPT-5.6 Sol's ARC-AGI-3 score—and what that tells us about evals
OpenAI's GPT-5.6 Sol jumped from 13.3% to 38.3% on ARC-AGI-3 by keeping reasoning in context and using compaction. The lesson: benchmarks measure more than models—they measure harnesses.