AI Matters
Why it matters
Anthropic's Opus 5 far outscores rivals on a tough new reasoning test
The Decoder Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence
Anthropic's Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly quadrupling the previous best score set by GPT-5.6 Sol.
Why it matters
A leap in reasoning ability could mean AI models get noticeably better at solving problems that stumped earlier versions.
Latest brief · Jul 26, 2026