Anthropic ships Claude Opus 5 at half Fable 5's price. The frontier race just shifted from capability to economics.
Anthropic released Claude Opus 5 on July 24 at the same $5/$25 per million token pricing as Opus 4.8, delivering near-Fable 5 intelligence at half the cost. The model leads on Frontier-Bench, GDPval-AA, ARC-AGI-3, OSWorld 2.0, and Zapier AutomationBench, while scoring its lowest misalignment rating on Anthropic's automated behavioral audit. The launch signals that the competitive axis in frontier AI has shifted from peak capability to per-task economics.
Context from: Anthropic | VentureBeat | Digitalapplied
The decision it puts on your desk
Re-baseline your model routing this week. If you are paying Fable 5 prices for engineering and operations workloads, test Opus 5 on your own evals before July 31. The price gap is 50 percent. The benchmark gap on general coding and agentic tasks is inside noise. For specialist legal, clinical, or offensive-security workloads, the frontier tier still buys you something. For everything else, the economics have shifted. The labs are no longer competing on what their best model can do. They are competing on what your default model should cost.
Anthropic released Claude Opus 5 on Thursday at $5 per million input tokens and $25 per million output. Identical to Opus 4.8. Half of Fable 5.
And in nearly every benchmark that matters for engineering teams, it either leads or ties the frontier.

Anthropic now sells four tiers. Fable 5 sits at the top, priced at $10/$50 per million tokens, positioned for the longest autonomous jobs where the model has to stay coherent across hours or days with dense source material. Opus 5 is the new daily driver at $5/$25, the model you hand complex work to and review when it is done. Sonnet 5 is for work you run at scale where speed and cost per call decide what ships. Haiku 4.5 handles subagents and instant answers.
This is not a new capability ceiling. Anthropic did not claim Opus 5 beats Fable 5 outright. Fable 5 still wins on the longest, most autonomous jobs. The math is in the middle band. If most enterprise AI work happens there, and if Opus 5 is now the strongest model in that band at the lowest price of any frontier model, the front of the pack stops mattering for revenue.

The performance curves tell a different story than the scores
Anthropic's Frontier-Bench v0.1 chart plots every model's score against cost per task, with separate curves for the four effort settings: low, high, xhigh, and max. This is the chart that matters, and it changes what the headline numbers mean.

On Frontier-Bench v0.1, Opus 5 scored 43.3 percent at max effort. Opus 4.8 managed 21.1 percent at the same effort. Fable 5 reached 33.7 percent. GPT-5.6 Sol hit 34.4 percent.
Opus 5 leads the curve. At every effort setting, Opus 5 achieves higher scores at lower cost than any competing model. At max effort, it matches or exceeds Fable 5's peak score at roughly half the cost per task. The implication for anyone running agentic workloads at volume is not subtle. A workload that spends $4,000 a month on Fable 5 tokens costs $2,000 on Opus 5 at identical or better results, before any efficiency gain from fewer turns.
On CursorBench 3.2 at max effort, Anthropic puts Opus 5 within 0.5 percent of Fable 5's peak score at half the cost per task. Across the full chart, Opus 5 dominates every other model on the high, xhigh, and max effort settings.
The headline number is misleading if you read it in isolation. The real claim is the curve.
The knowledge work and problem-solving sweep
The performance story extends well beyond coding.

On ARC-AGI 3, a benchmark of novel problem-solving where models cannot rely on memorized patterns, Opus 5 scored 30.2 percent. Opus 4.8 scored 1.5 percent. GPT-5.6 Sol scored 7.8 percent. Anthropic describes the result as three times the next-best model's score. The gap between a model that fails almost every novel problem and one that solves a third of them is not a benchmark bump.
It is a category change.
On Zapier AutomationBench, which measures whether models can complete business workflows end to end, Opus 5's pass rate of 26 percent led all four models by nearly eight points. Previous models failed the churn-prevention workflow entirely. Opus 5 hit 100 percent, according to Zapier CEO Wade Foster.
On OSWorld 2.0, a computer-use benchmark, Opus 5 hit 70.6 percent against Fable 5's 66.1 percent at roughly a third of the cost. Even at its lowest effort setting, Opus 5 passes more tasks than any other model.
On BrowseComp, an agentic search benchmark, Opus 5 hit 90.8 percent against Fable 5's 87.4 percent and GPT-5.6 Sol's 90.4 percent. On Humanity's Last Exam without tools, Fable 5 edges Opus 5 by two-tenths of a point at 56.5 percent against 56.3 percent. With tools enabled, Opus 5 retakes the lead at 64.7 percent against 63.9 percent.
The four spots where Opus 5 does not win cluster in specialist domains. GPT-5.6 Sol leads DeepSWE v1.1 at 72.7 percent against Opus 5's 68.8 percent. The held-out Legal Agent Benchmark goes to Fable 5 at 13.3 percent against 11.7 percent, though every model fails most tasks here. Mythos 5 retains its lead on cybersecurity exploitation and biology research. On FrontierCode v1.1, Opus 5 and Fable 5 are a statistical tie at 53.4 percent against 53.5 percent.
The pattern is consistent. Opus 5 dominates general agentic and coding work. It concedes narrow ground on specialist domains and offensive security. The gap on coding benchmarks between Opus 5 and Fable 5 is inside noise.
The behavior upgrade is self-verification
The benchmarks tell half the story. The other half is what Opus 5 does when a benchmark fails to provide the tools it needs.
Given a drawing of a machine part with no way to view it, Opus 5 wrote its own computer vision pipeline to extract geometry from raw pixels. It solved the task repeatedly.
No competing model with the same constraints solved it in five attempts, Anthropic said.
Given a real bug in a popular open-source package manager, it found the root cause and fixed an edge case the community's patch had missed. A competing model patched the symptom and declared victory.
An engineer at a trading firm watched it build a market data feed from scratch and, finding no live feed to validate against, construct its own test harness to verify its parsing code.
"The evals where Opus 5 wins are bounded tasks with a specific outcome, which is where it's strongest," an Anthropic spokesperson said. "What those evals don't measure is duration. Opus 5 is the best tool for the jobs benchmarks can see, and Fable 5 is what you reach for when the job outruns the benchmark."
The hidden cost of enterprise AI is human review. Engineers checking the machine's work. A model that reliably checks its own work compresses that cost. Customers keep citing fewer turns, fewer passes, less time, not higher raw scores. The capability that matters for production is the model that verifies its output before handing it back.
The alignment chart is the underreported story
Anthropic's automated behavioral audit scores every recent model on overall misaligned behavior. The chart shows a clear stacking pattern.

Opus 5 scored 2.3. Opus 4.8 scored roughly 3.5. Sonnet 5 scored around 4. Fable 5 scored close to 5.
Anthropic says Opus 5 adheres to Claude's Constitution better than Opus 4.8, Sonnet 5, or Fable 5. It exhibits the lowest rates of deceptive behavior, the least susceptibility to being tricked into misuse, and is the safest model yet at avoiding reckless actions with hard-to-reverse side effects.
On dangerous capabilities, Anthropic's position is that Opus 5 does not advance the frontier. Evaluated alongside private-sector and government partners, it remains behind Mythos 5 in both biology research and offensive cybersecurity.
The OSS-Fuzz chart makes the design intent precise. Mythos 5 and Opus 5 identify software vulnerabilities with similar success, 80 percent versus 79.4 percent.
On exploit development, Mythos 5 solves 13 challenges. Opus 5 solves 4.

This asymmetry is by design. Anthropic deliberately avoided training Opus 5 on cyber tasks, yet it improved substantially anyway as a by-product of general capability gains. The model finds vulnerabilities at near-Mythos level. It remains far behind at turning those vulnerabilities into material cyber threats. The safeguards follow the same logic. Opus 5's cyber classifiers are proportionally less restrictive than Fable 5's, allowing vulnerability research in source code while blocking binary-based scanning, penetration testing, and exploit generation. Anthropic expects them to intervene around 85 percent less often than Fable 5's.
On biology the calculus runs the other way. Because Opus 5 carries a similar safeguard suite to Opus 4.8 rather than Fable 5's stricter one, it is now Anthropic's most capable generally available model for scientific research. Biology-related requests blocked on Fable 5 now route to Opus 5 rather than Opus 4.8.
What this means for the frontier competition
For three years, the model labs competed on what their best system could do on its best day. Benchmarks were a capability arms race.
The score that mattered was peak performance. Price was secondary.
Anthropic just flipped the script. Opus 5 is not the smartest model the company makes. Fable 5 and Mythos 5 hold that distinction.
But Opus 5 is the model Anthropic expects most enterprise customers to actually use, every day, for most tasks. It is the new default on Claude Max. It is the strongest model available on Claude Pro. A large population of paying users wakes up on it without changing a setting.
That is a market structure play, not a benchmark play. If most enterprise AI work happens in a middle band of difficulty where near-frontier intelligence at half the price beats frontier intelligence at full price, the front of the pack stops mattering for revenue.
The money is just behind it.
OpenAI faces the sharpest pressure. GPT-5.6 Sol trails Opus 5 on seven of the twelve evaluation families Anthropic published and leads only DeepSWE v1.1 by four points.
Google's Gemini 3.6 Flash and 3.5 Flash sit in a lower pricing tier but have not posted comparable agentic scores. DeepSeek V4 and Kimi K3 compete on open-weight economics but lack the self-verification behavior enterprise customers are now citing as the differentiator.
The efficiency numbers from early customers are directional but consistent. Harvey, the legal AI company, reported 26 percent fewer tokens for similar performance. A financial modeling team reported nine percentage points higher accuracy with a third fewer turns and 60 percent less time.
A trading firm reported one-seventh the reasoning tokens and under half the latency of Opus 4.8. These are self-reported. The direction matters more than the precision. Models that finish in fewer steps cost less to run, period.
Anthropic now carries a reported $380 billion valuation and roughly $9 billion in annualized revenue, with internal targets of $20 to $26 billion for 2026. That is underwritten by a $30 billion Azure compute deal alongside Google Cloud and Nvidia arrangements.
Holding Opus 5's price at Opus 4.8 levels while roughly doubling performance on key agentic benchmarks is a deliberate widening of the funnel. Every task that was marginal at Opus 4.8's cost-per-success becomes viable at Opus 5's. Every viable task is recurring token revenue.
The regulatory backdrop is hardening in parallel. A judge gave final approval this week to Anthropic's $1.5 billion copyright settlement with book authors. The US government moved in June to block foreign access to Anthropic's most advanced models. Which customers can buy which capabilities is now a policy question as much as a pricing question.
One thread is unresolved. Every benchmark figure comes from Anthropic's internal harness with Opus 4.8 as the fallback for safety-classifier refusals. The Frontier-Bench scores are not pure single-model results. If the fallback routing affects the numbers, the real-world gap may be narrower than the charts suggest. Production workloads will tell.
A second thread is sharper. The OSS-Fuzz chart shows Opus 5 at 79.4 percent vulnerability identification against Mythos 5 at 80 percent. That is a near-tie on a capability Anthropic explicitly did not train for. The general-capability gains that produced this near-tie will compound into the next release. By Opus 5.5 or Opus 6, the gap on vulnerability identification may close entirely. Anthropic's safeguard design assumes a gap that may not hold.
Source
← Previous
Corgi Hits $4B Valuation in Third Raise in Eight Weeks as AI Insurance Frenzy Defies Gravity
Next →
Klaimee Raises $5.5M Seed to Insure What Cyber and E&O Policies Won't Touch: Autonomous AI Agents
The Morning Recap
We sift the day. You make the calls.
The decision on every story, in your inbox before your first call.