Anthropic releases Claude Opus 5, reaching Fable 5-level coding scores at half the per-task cost
Anthropic released Claude Opus 5 on July 24, 2026, at the same price as its predecessor Opus 4.8: $5 per million input tokens and $25 per million output tokens per Anthropic's announcement. On OSWorld 2.0, the model surpassed Fable 5's best computer-use score at just over one-third of Fable 5's per-task cost. On CursorBench 3.2, Opus 5 came within 0.5% of Fable 5's peak coding score at max effort while costing half as much per task.
What happened
On Frontier-Bench v0.1, Opus 5 topped all other models and more than doubled Opus 4.8's benchmark performance at a lower cost per task, per Anthropic's release. The model also set new state-of-the-art results on GDPval-AA, an Artificial Analysis knowledge-work evaluation. On ARC-AGI 3, Anthropic reported Opus 5's score as three times that of the next-best model. On Zapier AutomationBench, which tests end-to-end business-task completion, Opus 5's pass rate ran at roughly 1.5 times the next-best model for the same cost per task.
One Frontier-Bench task captures the capability improvement concretely: Opus 5 was given a drawing of a machine part and asked to produce a 3D FreeCAD reconstruction, but was deliberately given no way to view the drawing directly. The model wrote its own computer vision pipeline, pulled geometry from raw pixels, and rebuilt the machine part successfully across repeated attempts. No competing model with the same setup solved it after five tries, per Anthropic.
Life sciences results improved across the board as well. On organic chemistry tasks such as inferring molecular structures from spectroscopy data, Opus 5 scored 10.2 percentage points higher than Opus 4.8. On protein-related tasks, including predicting how sequence variations affect protein function, it scored 7.7 percentage points higher.
The model launched across Claude.ai, the Claude API (model ID: claude-opus-5), Claude Code, and Claude Cowork. It became the new default on Claude Max and the strongest model on Claude Pro. Consistent with prior Opus releases, Opus 5 carries no data retention requirements for general access; Fable 5 had imposed a 30-day retention period that blocked deployment in regulated industries. A Fast mode runs at approximately 2.5 times the default speed, available at twice the base API price.
Two beta features shipped alongside the model: mid-conversation tool changes, which let developers modify a conversation's available tools without resetting the prompt cache; and automatic API fallbacks, which redirect requests flagged by safety classifiers to another model instead of returning a block.
Why it matters
For teams running production agents on Fable 5, the operative change is cost. Opus 5 reaches Fable 5's coding-benchmark ceiling at roughly half the per-task price on CursorBench and beats Fable 5 outright on computer use at one-third the cost. The no-retention policy extends Opus 5's reach into healthcare, finance, and legal pipelines where Fable 5's 30-day requirement was a disqualifying constraint, without any architectural workaround.
Alignment improved in parallel. Anthropic's automated behavioral audit scored Opus 5 at 2.3 on overall misaligned behavior, the lowest result among its recent models per Anthropic's announcement. The company described it as its most aligned model to date, with lower rates of deceptive behavior and reduced susceptibility to misuse-triggering prompts compared to Sonnet 5, Opus 4.8, and Fable 5.
On safety evaluations conducted with private-sector and government partners, Opus 5 did not advance the frontier on biological research or offensive cybersecurity. It remained behind Mythos 5 on exploit development even though its vulnerability-identification capability now approaches Mythos 5's level. The cyber classifiers on Opus 5 are less restrictive than those on Fable 5: they allow source-code vulnerability scanning but block binary-based scanning, penetration testing, and exploit generation for users outside the Cyber Verification Program.
Context and reactions
Early-access customers cited gains across production workloads. Wade Foster, CEO of Zapier, said in Anthropic's announcement that Opus 5 topped AutomationBench without spending more tokens than prior Claude models and ran a complete churn-prevention workflow from a raw account-health workbook, end to end, while "previous models didn't pass." Scott Wu, CEO of Cognition, the company behind the Devin agent platform, said Opus 5 "approaches Fable-level performance at half the cost" on FrontierCode 1.1, with particular strength in debugging and root-cause analysis.
Box reported Opus 5 outperforms Opus 4.8 by 8% on its internal benchmark, with gains of 11% on data analysis workflows and 17% on due diligence tasks per Anthropic's announcement. A legal-tech early-access partner said Opus 5 achieved similar accuracy to Opus 4.8 while generating 26% fewer tokens on average at max reasoning, which cuts per-query costs further than the base pricing change alone.
Enterprises enrolled in the Cyber Verification Program received immediate access to a version of Opus 5 with fewer security restrictions. For other users, CVP enrollment remains the pathway to binary-based scanning and related capabilities. Requests that trigger cyber classifiers in Claude.ai, Claude Code, and Claude Cowork fall back to Opus 4.8 by default.
What to watch next
Third-party benchmark replications from Artificial Analysis and LiveBench will be the first independent check on Anthropic's Frontier-Bench v0.1 and ARC-AGI 3 figures, which come from Anthropic's own test runs. Anthropic has not announced a timeline for lifting penetration-testing or exploit-generation restrictions for the broader Opus 5 user base.
Sources
- Introducing Claude Opus 5: Anthropic, July 24, 2026
- Anthropic launches Opus 5: TechCrunch, July 24, 2026
- Claude Opus 5 matches Fable on coding at half the price: The Next Web, July 24, 2026
