AI Matters
Why it matters
Anthropic's Opus 5 sets a big lead on a key reasoning test
The Decoder Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence
Anthropic's Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly four times GPT-5.6 Sol's previous best score of 7.8 percent.
Why it matters
The model solved reasoning puzzles no prior AI could, a sign general problem-solving is advancing fast.
Latest brief · Jul 26, 2026