Claude Opus 5 brings near-frontier performance at half the price: $5 per million input tokens, $25 per million output
Claude Opus 5 is Anthropic's newest Opus model, released on July 24, 2026 and priced at $5 per million input tokens and $25 per million output tokens. Anthropic positions it as a model that comes close to the frontier intelligence of Claude Fable 5 at half the price. On the CursorBench 3.2 coding benchmark it lands within 0.5% of Fable 5's peak score, and on ARC-AGI 3 it scores three times as high as the next-best model. ASAP summarizes the numbers and conditions of this release from Anthropic's own announcement and the AWS engineering blog.
The price tag is the real headline of this release
The pricing, not the leaderboard position, is the biggest variable in the Opus 5 release: $5 per million input tokens and $25 per million output tokens. Anthropic frames nearly every performance claim in terms of cost per task rather than absolute score. On CursorBench 3.2 the model reaches within 0.5% of Fable 5's peak at half the cost, and on OSWorld 2.0 it surpasses Fable 5's result at roughly one-third the cost. A separate Fast Mode runs at twice the base price for roughly 2.5 times the speed.
This structure signals a shift in how models are sold. A launch that leads with a benchmark win speaks the language of research competition, while a launch that leads with cost per task speaks to whoever controls the operating budget. Listing Fast Mode explicitly at 2x price for 2.5x speed separates latency from capability and turns it into a purchasable resource. Teams can attach the fast tier to interactive work and the base tier to overnight batch jobs, and that choice is built into the price sheet from day one.
Benchmarks are presented as cost-adjusted performance, not raw scores
Frontier-Bench v0.1 results are the clearest example: Opus 5 surpasses all other models and more than doubles Opus 4.8's performance at a lower cost per task. On ARC-AGI 3 it scores three times as high as the next-best model, and its Zapier AutomationBench pass rate is around 1.5x the next-best model for the same cost per task. Those three benchmarks target reasoning, generalization, and practical automation respectively, which places the launch well outside a single coding metric.
One detail deserves care: these figures are mostly relative expressions, given as multiples and percentage points. Without published absolute scores, the same "three times" means very different things depending on which model was next-best and how low the baseline sat. A 3x gain over a weak baseline and a 3x gain over a strong one are not the same thing in production. Frontier-Bench v0.1 is also, as its version number indicates, an early-stage benchmark, so there is not yet a basis for treating it as an industry-standard yardstick. The numbers themselves come straight from Anthropic's announcement; the open question lives in the baselines that were not disclosed.
Why the life-sciences numbers were broken out separately
The life-sciences results are reported only as gains over the previous generation: 10.2 percentage points higher than Opus 4.8 on organic chemistry and 7.7 percentage points higher on protein function prediction. Breaking these two items out of an announcement otherwise centered on coding and agentic work is a deliberate editorial choice.
That placement reads as a market signal. Coding performance is a crowded field where several models trade the lead, while chemistry and biology tasks carry high verification costs and few published competitor numbers. Reporting these two items purely as generation-over-generation improvement, rather than against rivals, is the conventional way to claim ground in a category where no ranking has formed yet. For pharmaceutical and biotech research groups in Korea, these two figures work better as a hypothesis to reproduce on internal tasks than as a basis for a procurement decision.
What the 2.3 safety score does and does not tell you
Anthropic's automated behavioral audit score for Opus 5 is 2.3, which the company describes as the lowest misaligned behavior among its recent models. The announcement also states plainly that the model remains behind Mythos 5 on cybersecurity exploitation tasks, and that safeguards stay similar to those on Opus 4.8.
The operational detail in the AWS documentation is more interesting. In higher-risk cybersecurity areas, Opus 5 may fall back to Opus 4.8, with the user notified when it happens. That design constrains capability through product-level routing rather than by degrading the model during training. The approach is sensible because it avoids trading capability against safety wholesale, but it also means an identical API call can be served by a different model depending on context. Organizations that run reproducible evaluations or keep audit logs need to establish in advance which request types are subject to fallback. A 2.3 is the output of one specific behavioral-audit instrument, and it does not represent deployment risk in general.
What Korean teams should calculate before adopting it
Availability spans Claude.ai, Claude Code, Claude Cowork, and the Claude API, and on Amazon Bedrock the model is invoked as global.anthropic.claude-opus-5. AWS lists four initial regions: US East (N. Virginia), Asia Pacific (Melbourne), Europe (Ireland), and Europe (Stockholm), and points to the Bedrock documentation for the full list.
Latency and data residency are the first variables a Korean organization needs to work through. Seoul is not among the initial regions AWS named, so which region a Korean request transits affects both response time and where data is processed. AWS also notes support for adding and removing tools dynamically mid-conversation, and for teams designing long-running agents that capability matters as much as the price sheet. Being able to change the tool list mid-task reduces the need to end and restart a session, which removes the repeated cost of re-sending context.
What remains unanswered
Anthropic's announcement does not include a context-window figure, and it leaves most absolute benchmark scores and comparison models unnamed. What can be established today therefore stops at pricing, platform availability, and the direction of the improvements; actual cost-performance on a given workload requires measurement against each organization's own data. With Anthropic shipping its fourth model in two months, the pace itself is worth weighing, because pinning a pipeline tightly to one version carries its own cost.
Source: ASAP summary based on Anthropic's announcement "Introducing Claude Opus 5" (July 24, 2026) and the AWS Machine Learning Blog post "Introducing Claude Opus 5 on AWS: Anthropic's most capable Opus model" (July 24, 2026).

AI & tech,
read in depth
Beyond the headlines — into the context and the structure
AGI Soon As Possible · asapai.co.kr