AI Matters

Why it matters

Anthropic's Opus 5 far outscores rivals on a tough new reasoning test

The Decoder Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence

Anthropic's Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly quadrupling the previous best score set by GPT-5.6 Sol.

Why it matters

A leap in reasoning ability could mean AI models get noticeably better at solving problems that stumped earlier versions.