AI Matters

Why it matters

OpenAI claims new model beats Anthropic's on reasoning test, with caveats

The Decoder OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings

OpenAI said GPT-5.6 Sol scored higher than Anthropic's Opus 5 on the ARC-AGI-3 benchmark, but only when using its own API features rather than the standard test setup.

Why it matters

The claimed win only holds using OpenAI's own API settings, showing how much test conditions can shape AI benchmark results.