ai, engineering, llm,

Claude Opus 5: State-of-the-Art at Half the Price

Cui Cui Follow Jul 27, 2026 · 5 mins read
Claude Opus 5: State-of-the-Art at Half the Price
Share this

Anthropic shipped Claude Opus 5 on July 24 and it lands with a clear thesis: frontier-class intelligence doesn’t require frontier-class spend. Opus 5 sits below Fable 5 on cybersecurity-specific evals, but it tops every other model on software engineering, knowledge work, and agentic tasks — at roughly half the cost per task.

If you run coding agents, research pipelines, or any workflow where you’re currently paying Fable 5 prices, this is worth your attention.


The Value Equation

Anthropic built Opus 5 around a variable effort setting that lets you tune intelligence versus token consumption. At max effort, Opus 5 approaches Fable 5’s performance. At low effort, it still outperforms most other models. This isn’t a trimmed-down model with a cost discount — it’s a model that surfaces the right level of reasoning for the task.

The headline numbers:

  • CursorBench 3.2 — within 0.5% of Fable 5’s peak at max effort, but at half the cost per task
  • Frontier-Bench v0.1 — tops all models, more than doubles Opus 4.8’s score at lower per-task cost
  • ARC-AGI 3 — 3× the score of the next-best model on novel problem-solving
  • Zapier AutomationBench — 1.5× the pass rate of next-best at the same token spend; even its lowest-effort setting outperforms every other model
  • OSWorld 2.0 — best-in-class computer use performance at any cost level, beating Fable 5’s peak result at roughly one-third the cost

That last point deserves a pause. Beating Fable 5 on computer use while spending 67% less per task isn’t a marginal improvement — it reshapes what you should be paying for.


What “Agentic” Actually Looks Like Now

The benchmark numbers are one thing. The qualitative examples Anthropic shared are what actually tell you where the capability ceiling sits.

The FreeCAD case. Frontier-Bench gave Opus 5 a drawing of a machine part and asked it to reconstruct it as a 3D model — then deliberately gave the model no direct way to view the drawing. Opus 5 responded by writing its own computer vision pipeline to extract geometry from raw pixels, then rebuilt the full part. It solved this repeatedly. No other tested model could do it once in five attempts.

The package manager bug. Given a real bug in a popular open-source package manager, Opus 5 found the root cause and fixed an edge case that the existing community patch had missed. A competing model fixed only the surface symptom and reported success.

The trading firm case. An engineer used Opus 5 to build a complete market data feed for a new exchange in a single session — a task no previous model could finish even with extensive upfront planning. Finding no live feed to validate against, the model built its own test harness to check that its output parsed the exchange’s data correctly.

These aren’t cherry-picked prompts. They’re engineering tasks from real users, and the through-line is the same: Opus 5 verifies its own work, recovers from blocked paths, and iterates until it actually succeeds. Earlier models tend to converge on a plausible-looking output and stop.


Science and Research Improvements

Beyond software engineering, Opus 5 makes meaningful gains on scientific benchmarks:

  • Outperforms Opus 4.8 on every life sciences evaluation Anthropic runs
  • +10.2 percentage points on organic chemistry tasks (inferring molecular structures from spectroscopy data)
  • +7.7 percentage points on protein variant function prediction

A team doing genomics analysis described Opus 5 as behaving “more like a careful scientist than any model we’ve run” — reaching for the right statistical tests, cross-checking results by independent methods, and staying coherent across multi-step analyses.

For AI engineering teams working in bioinformatics, drug discovery, or any domain requiring both coding and scientific reasoning, this combination is relevant.


What This Changes in Practice

The practical implication of Opus 5 depends on where you currently sit:

If you’re on Opus 4.8: This is an unambiguous upgrade. More capable across every eval Anthropic publishes, same cost profile. The consistency improvement — mentioned across multiple early-access customers — is particularly worth noting for production pipelines where variance is expensive.

If you’re on Fable 5: The calculus depends on your workload. On cybersecurity-specific tasks, Fable 5 still leads. On most other categories — especially coding, agentic tasks, and knowledge work — Opus 5 delivers near-equivalent results at roughly half the per-task cost. For high-volume pipelines, that’s a meaningful budget unlock.

If you’re running Claude Code: Devin reports Opus 5 as particularly strong on debugging and root-cause analysis — the tasks where coding agents typically break down. Cursor reports it’s within 0.5% of Fable 5 on their benchmark, with the same behavioral characteristics. Lovable calls it steadier run-to-run than prior models, which matters more for a builder platform than peak benchmark scores.


Availability

Opus 5 is now the default model on Claude Max and the strongest available on Claude Pro. API access is available through Anthropic’s API, Amazon Bedrock, and Google Cloud Vertex AI.

The effort parameter is configurable through the API, letting you set it at the task level — meaning you can run low-effort passes for triage and high-effort passes for complex synthesis in the same pipeline.


Bottom Line

Anthropic’s message with Opus 5 is clear: the cost-intelligence tradeoff curve just moved. A model that tops the field on software engineering and agentic tasks while running at half the cost of the frontier isn’t a category compromise — it’s a category definition.

For engineers running agentic systems, the interesting question now is whether the remaining gap between Opus 5 and Fable 5 matters for your specific workload. If your tasks live in the coding, automation, and knowledge-work space that Opus 5 was built for, the answer is probably no.

That’s a significant shift in where the default should land.

Join Newsletter
Get the latest news right in your inbox. We never spam!
Cui
Written by Cui Follow
Hi, I am Z, the coder for cuizhanming.com!

Click to load Disqus comments