Claude Opus 5 went live on July 24 as the new default model on Claude Max, priced at the same $5/$25 per million tokens as its predecessor. A day later, independent benchmarking from Artificial Analysis put it narrowly ahead of Fable 5 on its composite Intelligence Index, and comfortably ahead on tests built around real agentic work.
Opus 5 Narrowly Overtakes Fable 5 for the Top Spot on the Intelligence Index
Anthropic's own announcement is deliberately modest. It describes Opus 5 as a model that comes close to the frontier intelligence of Claude Fable 5, at half Fable 5's token price — $5 and $25 per million input and output tokens, against Fable 5's $10 and $50. Artificial Analysis's independent scoring tells a slightly different story. On the Artificial Analysis Intelligence Index, a composite drawn from nine evaluations including GDPval-AA v2, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, and GPQA Diamond, Opus 5 at max effort scores 61, one point ahead of Fable 5's 60. GPT-5.6 Sol's three-tier launch put OpenAI's model one point behind Fable 5 at 59, and Kimi K3, Moonshot AI's open-weight entrant, follows two points further back at 57. Opus 4.8, the model Opus 5 replaces, trails the new release at 56.
The margin is close enough that Artificial Analysis's own coverage treats the top two models as statistically equivalent at the front of the index. But it is notable that Anthropic's second-tier model has out-scored its own flagship sibling on an independently run composite this soon after Fable 5's June debut, and the identities of the closest competitors track what this publication has already reported: a one-point gap behind Fable 5 for GPT-5.6 Sol, and a further gap for Kimi K3 as a strong but still-trailing open-weight model.
On Agentic Knowledge-Work Benchmarks, Opus 5's Lead Widens Well Beyond One Point
Anthropic's launch page leans harder on agentic, real-world work than on the composite index, and that is where Artificial Analysis's numbers show the widest separation. On AA-Briefcase, Artificial Analysis's own benchmark for agentic knowledge work, built on its open-source Stirrup reference agent harness, Opus 5 at max effort scores an Elo of 1,720, 146 points ahead of Fable 5's 1,574. On GDPval-AA v2, a separate benchmark of agentic professional tasks, Opus 5 reaches 1,861 Elo, more than 100 points ahead of both Fable 5 and GPT-5.6 Sol, and more than 114 points ahead of Fable 5 specifically — Artificial Analysis has not published Fable 5's exact GDPval-AA v2 score, so that gap is a documented lower bound rather than a precise figure.
Anthropic's own customer testimonials point in the same direction, though these are vendor-reported and not yet independently reproduced. Zapier's chief executive said Opus 5 completed a full, multi-step churn-prevention workflow on the company's AutomationBench without exceeding the token spend of earlier Claude models, where prior models had not passed at all. Cursor's team said the model landed within half a percentage point of Fable 5's peak coding score on CursorBench 3.2, at roughly half the cost per task. Independent coding-leaderboard results, separate from the figures below, were not yet available at the time of publication.
Opus 5 Costs More Per Task Than Sonnet 5 and Opus 4.8, But Less Than Fable 5
The "half the price" framing refers to Opus 5's published token rates against Fable 5's, not to what a task actually costs once a model's verbosity and reasoning effort are factored in. Artificial Analysis's task-weighted accounting, which measures how many tokens each model actually spends completing an Intelligence Index task rather than just the price per million, puts Opus 5 at max effort at $2.03 per task — a 26% reduction from Fable 5's $2.75, but still above Opus 4.8's $1.80 and above Claude Sonnet 5's $1.53. The two numbers, 50% and 26%, are measuring different things: one is a list-price ratio, the other is a real usage-weighted cost, and Opus 5 is more verbose per task than its token pricing alone would suggest.
The Claude lineup now forms a clean three-tier price ladder: Sonnet 5 at $2 and $10 per million input and output tokens for routine work, Opus 5 at $5 and $25 for complex semi-autonomous tasks, and Fable 5 at $10 and $50 for jobs where cost is not the deciding factor. Opus 5 is now the default model on Claude Max and the strongest model available on Claude Pro, so a meaningful share of subscriber traffic moves onto it without any setting changes. A Fast mode is available at roughly 2.5 times the default speed for twice Opus 5's base price, and unlike Fable 5, Opus 5 carries no mandatory data-retention requirement for general access.
Anthropic Keeps Opus 5 Behind Mythos 5 on Cyber Exploit Development, Not Vulnerability Discovery
Opus 5's safety posture is the one place Anthropic's own copy draws a firm line rather than a comparative claim: the model, in Anthropic's own words, comes close to the frontier intelligence of Fable 5, but the company is explicit that it stays behind Mythos 5 on cybersecurity tasks specifically. The company's system card breaks that gap down by phase. On OSS-Fuzz, an internal evaluation covering both vulnerability discovery and exploit development, Opus 5 finds vulnerabilities at a rate Anthropic describes as close to Mythos 5's, but its ability to turn a discovered vulnerability into a working exploit lags well behind Mythos 5's. That distinction matters for anyone assessing Opus 5 for defensive security work: finding a flaw and weaponizing it are different skills, and Anthropic is only claiming parity on the former.
Because of that gap, Opus 5's cyber safety classifiers block binary-based vulnerability scanning, penetration testing, and exploit generation outright, while still allowing source-code vulnerability discovery. Anthropic expects those classifiers to intervene roughly 85% less often than Fable 5's, with flagged requests falling back to Opus 4.8 by default. The company also reports Opus 5 as its best-scoring model yet on an internal alignment audit, at 2.3 on overall misaligned behavior, lower than the scores it has recorded for Opus 4.8, Sonnet 5, or Fable 5. That narrower cyber posture arrives against a backdrop of rising disclosure activity elsewhere in the Claude lineup: the surge in disclosures tied to Mythos 5 reported a threefold jump in reported vulnerabilities following that model's release, a pattern that makes Opus 5's tighter restrictions look like a deliberate trade-off rather than an oversight.
Taken together, the two evaluations point to a model Anthropic is comfortable putting in front of the largest share of its subscriber base — as the Claude Max default and the strongest Claude Pro option — while still keeping meaningfully more restrictive guardrails around it than a general jump in benchmark scores alone would suggest.





Comments (0)
Please sign in to join the discussion.
No comments yet.
Be the first to share your perspective on this topic.