AGI Soon As Possible · Deep reads on AI & tech
Article

Claude Fable 5.1 Cuts Only Its Cache Read Price by 75% and Lowers Total Cost by up to 45%

2026-09-02 · 14 min read

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026, stating that the two are the same model with different levels of safeguards. Input at $10 per million tokens and output at $50 per million tokens are unchanged from Fable 5, and the single item that changed is cache reads, now $0.25 per million tokens, a 75% reduction. Anthropic says that change alone reduces typical workload costs by roughly 25% and highly agentic workload costs by up to roughly 45%, since cache reads make up most of the bill in context-heavy, tool-heavy work. On the agentic science benchmark Terminal-Bench-Science 0.1, Fable 5.1 scores 52.6% against 24.7% for Fable 5 and 29.0% for Opus 5, and Mythos 5.1 reaches 60.9% on Terminal-Bench 4.0.

One Line Item, Cache Reads, Moves the Entire Price Sheet

The cost reduction in Claude Fable 5.1 comes entirely from one line item, cache reads dropping 75% to $0.25 per million tokens. Input and output rates remain identical to Fable 5 at $10 and $50 per million tokens. In the indexed cost chart Anthropic published, with Fable 5 set to 100, Fable 5.1 lands at 75 for typical workloads and 55 for highly agentic ones.

Cache reads are the portion where the model rereads inputs it has already processed and stored. The more a workload references the same context repeatedly, the larger that portion grows in the bill. Anthropic defines highly agentic work as context-heavy, tool-heavy work where cache reads make up most of the cost, which is exactly the profile that benefits here.

What is easy to miss is that the reduction is not uniform. The 25% and 45% figures are not posted discount rates but observed outcomes for specific usage patterns. Anthropic states they were measured at default effort over four weeks of actual usage in August 2026, with typical workload covering Fable usage across Claude Enterprise, Claude Code, and the API. A team running short one-shot calls with little caching sees close to nothing from this change, while a team running long-context coding agents for hours approaches the 45% figure.

The character of the pricing move also changed. Rather than cutting the same percentage for everyone, Anthropic cut the cost of one usage shape, and that shape happens to match the long-running autonomous work the rest of the announcement emphasizes. The price change and the product direction point at the same place.

Only a Few Gaps in the Benchmark Table Clear the Noise Floor

On Terminal-Bench-Science 0.1, Fable 5.1 scores 52.6% against 24.7% for Fable 5, with Opus 5 at 29.0% and GPT-5.6 Sol at 22.4%. Elsewhere in the table, Terminal-Bench 4.0 shows 55.8% for Fable 5.1 and 60.9% for Mythos 5.1, against 42.0% for Fable 5 and 52.3% for Opus 5. On AutomationBench, a business workflow benchmark, Fable 5.1 reaches 31.4% against 17.1% for Fable 5. The knowledge work metric GDPval-AA v2 reads 1853 against 1723 for Fable 5 and 1824 for Opus 5.

Computer use scores on OSWorld 2.0 are 77.9% partial and 41.7% strict, with no GPT-5.6 Sol figure shown. Humanity's Last Exam reads 60.9% without tools and 65.0% with tools. On CursorBench 3.2.0, Fable 5.1 scores 73.4% against 70.5% for Fable 5, 70.0% for Opus 5, and 67.2% for GPT-5.6 Sol.

Two footnotes decide how this table should be read. First, the standard error on Terminal-Bench-Science 0.1 is ±3.5 to 4.5 points per model. Anthropic itself notes that the public leaderboard reports Opus 5 at 30.0% and Fable 5 at 21.4%, while its own setup reproduces them at 29.0% and 24.7%, both within noise. A 27.9-point gap like 52.6% against 24.7% clears that margin easily, but a gap under 3 points like 73.4% against 70.5% on CursorBench does not carry the same weight.

Second, the scoring conditions were set against the Fable models. Fable 5.1 was evaluated with production safeguards enabled, and on tasks where those safeguards intervened, Fable 5.1 and Fable 5 scored zero on OSWorld 2.0, with Fable 5 also scoring zero on AutomationBench. In other intervention cases, cybersecurity tasks were completed by Claude Opus 4.8 and biology tasks by Claude Opus 5. Anthropic states this likely reduced the performance of both Fable models. Read alongside competitors evaluated without those constraints, this table is a conservative reading of the Fable line.

The Safeguards Were Not Loosened, the Routing Was Changed

The concept behind Anthropic's safeguard changes is precision rather than relaxation. In cybersecurity, the newest safeguards block 60% fewer false positives than before, and Claude Code users can expect roughly 60% fewer safeguard interventions per session. In biology, safeguards fire 85% less often for benign elementary biology and medical questions relative to those that shipped with Fable 5.

The substance of the change is that Fable 5.1 can now be used to identify software vulnerabilities in source code. Developing exploits for those vulnerabilities remains off limits. Dual-use tasks including penetration testing, exploit generation, and binary-based vulnerability scanning are still redirected to Anthropic's Opus models, as are life sciences research and development queries.

Structurally, this is no longer a question of whether a request is blocked but of which model handles it. High-risk requests are routed to a differently capable model rather than refused, so a user receives another model's answer instead of a refusal notice. That reduces friction for defensive security teams, though the announcement does not say whether users can always tell when a request has been rerouted.

The Mythos 5.1 evaluation results belong next to this. Anthropic ran cyber evaluations with cybersecurity safeguards off and reports that Mythos 5.1 demonstrates the strongest cyber capabilities of any model it has released while still falling within the lower risk category of its Frontier Compliance Framework. On chemical and biological risk, capabilities exceed those of Mythos 5 but fall short of the next tier defined in the Responsible Scaling Policy, so Mythos 5.1 ships with the same safeguards applied to Mythos 5.

A 50% Hit Rate and a Venus Map Are Compressed Engineering, Not Discovery

Protein binders designed by Mythos 5.1 reached a hit rate of nearly 50% across 12 targets, against the 10 to 15% Anthropic describes as typical in protein design today. On three targets, binding affinities were 10 times higher than the best designs submitted to Adaptyv Bio's protein design competitions. Anthropic gave the model access to open-source protein design and folding tools and sent its designs to two external organizations for experimental validation.

A footnote sets the boundary of that claim. The three targets are EGFR, Nipah G, and 15-PGDH, and the Nipah G comparison is against de novo designs targeting the receptor-binding site on the G head. A separate entry from Nick Boyd and Escalante Bio targeting the stalk reached roughly 1.4 nM, comparable to Anthropic's best binder, as Anthropic states itself. The 10-times figure does not hold across every comparison.

The Venus result has a different character. Fable 5.1 trained a neural network on radar images taken by NASA's Magellan mission more than 30 years ago, plus an existing map covering a fifth of the planet, to produce a new high-resolution elevation map of a third of Venus. Detail resolves down to 2 to 3 kilometers rather than 10 to 20, and heights are up to 25% more accurate. Anthropic is releasing the map under a Creative Commons license ahead of NASA's VERITAS and ESA's EnVision missions.

The computational biology case is the most practical of the three. Mythos 5.1 wrote custom GPU kernels and cached intermediate results to speed up seven open-source deep learning models by up to 2.5 times with identical outputs. On genome-wide analyses, estimated GPU costs fell 30 to 60%. Anthropic notes such optimization normally takes a team of performance engineers weeks and is often unaffordable for academic labs, while Mythos 5.1 did it in days using publicly available source code alone.

Placed side by side, the three share a structure. None established a new scientific theory; all compressed engineering work that used to consume specialist time. Training a network to build a map, rewriting GPU kernels, and iterating design tools to filter candidates are tasks where the method was already known and skilled labor was the bottleneck. Even taking Anthropic's framing at face value, what is demonstrated here is not discovery itself but a drop in the unit cost of the pipeline leading to it.

The Quietest Change Is the Anti-Distillation Clause

The anti-distillation measure shipping with Fable 5.1 makes it impossible for API accounts created from launch day onward to manually edit Claude's prior context in a multi-turn conversation while preserving the transcript of Claude's prior thinking. Anthropic describes this as a common, publicly documented distillation technique that let distillers illicitly extract Claude's thinking. Distillation is often run at industrial scale using thousands of fake accounts, and Anthropic frames it as a safety risk because extracted capabilities can then be released without adequate safeguards.

The rollout is staggered. Existing accounts are not currently affected, but the change applies to all users with future model releases, at which point a small number of customers' custom integrations will break. This is effectively the only item in the entire announcement that may require a developer to change code, and it sits quietly between the benchmark table and the science section.

Any team that assembles multi-turn context by hand before passing it to the API should check this now. That is especially true for setups that summarize or trim conversation history while retaining thinking transcripts, since behavior will change when the team moves to a new account or a future model.

The Launch Answers Price, Data Retention, and False Positives at Once

Anthropic states outright that Fable 5.1 takes important steps toward addressing customer feedback on price, data retention, and safeguards. The data retention answer is Enterprise Frontier Safeguards, or EFS. Customer data is stored in cloud infrastructure controlled entirely by the customer rather than by Anthropic, and any human review is by default performed by the customer, which Anthropic says delivers the privacy of a zero data retention policy while remaining state of the art at preventing adversarial use.

EFS was developed with more than 100 customers across financial services, healthcare, manufacturing, telecom, law, retail, and the public sector, together with cloud partners at Amazon Web Services, Google Cloud, and Microsoft Azure. It will be supported on Claude Code, Claude Enterprise, the Claude Platform, Amazon Bedrock, Claude Platform on AWS, Google's Agent Platform, and Microsoft Foundry, rolling out in phases starting this fall. Until then, eligible customers can use Fable 5.1 and Fable 5 with zero data retention.

The same document carries the EU AI Act response. In July 2026 Anthropic signed the EU AI Act's Code of Practice on Transparency of AI-Generated Content alongside 190 other signatories, which required a watermark on outputs of models released after August 2, 2026. The watermark is a numerical way of determining the likelihood that Claude was involved in writing a text, is invisible to anyone without the detection API, and contains no information about the user, their organization, or their conversations. The detection API is in private preview for eligible organizations such as regulators, law enforcement, media, fact-checkers, independent researchers, educational organizations, and EU civil society groups.

Taken together, the actual subject of the announcement becomes visible. Benchmark scores are laid out in a table, but by volume and placement the center of gravity is the removal of enterprise adoption objections one by one. Too expensive is answered by the cache read cut, data leaves our perimeter by EFS, and legitimate work keeps getting blocked by the 60% and 85% reductions in safeguard firing. It reads less like a model announcement than an adoption-barrier announcement.

Three Things Teams Outside the US Should Calculate Now

The first item to check is the share of your own bill that cache reads represent. The 25% and 45% figures were measured on typical and highly agentic workloads respectively, so a workload dominated by short one-shot calls sees a far smaller reduction. Input at $10 and output at $50 per million tokens are unchanged, so the calculation starts from cache read volume.

The second is access to Mythos 5.1. It is currently available only to a set of US organizations, through the Cyber Verification Program for vetted defensive security work and the Life Sciences Verification Program built in partnership with the US government, which has enrolled its first participants. Anthropic says it is coordinating with the US government to expand access to a broader set of domestic and international partners as quickly as possible, but no terms or timing for organizations outside the US appear in the announcement. Security teams that need penetration testing or exploit generation will keep reaching those capabilities only through the Opus routing path.

The third is the scope of EFS and the watermark detection API. EFS rolls out in phases starting this fall, and the detection API is open mainly to organizations carrying obligations under EU law. Most companies outside those categories are not in scope for either right now, so a team with data retention requirements should first confirm whether it qualifies for zero data retention during the EFS interim period.

As of September 1, 2026, the confirmed facts are that Fable 5.1 and Mythos 5.1 are the same model separated only by safeguards, that cache reads dropped 75% to $0.25 per million tokens for estimated savings of 25% and up to 45%, the benchmark figures including 52.6% on Terminal-Bench-Science 0.1, a 60% reduction in cyber false positives and an 85% reduction in biology safeguard firing on benign queries, and Mythos 5.1's roughly 50% binder hit rate and the Venus map resolving 2 to 3 kilometers. Availability timing for Mythos 5.1 outside the US, the phase schedule for EFS, and the expansion scope of the detection API are not disclosed.

Source: ASAP summary based on Anthropic's official announcement "Introducing Claude Fable 5.1 and Claude Mythos 5.1" (September 1, 2026). Cited facts include the same-model-different-safeguards framing and trusted-access-only availability of Mythos 5.1, unchanged input and output pricing at $10 and $50 per million tokens with cache reads cut 75% to $0.25, estimated savings of roughly 25% on typical workloads and up to roughly 45% on highly agentic ones with indexed costs of 75 and 55 against Fable 5 at 100 measured over four weeks of August 2026 usage at default effort, Terminal-Bench-Science 0.1 at 52.6% against 24.7%, 29.0%, and 22.4% with a standard error of ±3.5 to 4.5 points and the public leaderboard reproduction note, Terminal-Bench 4.0 at 55.8% with Mythos 5.1 at 60.9% against 42.0%, 52.3%, and 37.3%, GDPval-AA v2 at 1853 against 1723, 1824, and 1711, OSWorld 2.0 at 77.9% partial and 41.7% strict, Humanity's Last Exam at 60.9% without tools and 65.0% with tools, AutomationBench at 31.4% against 17.1%, 26.9%, and 19.6%, CursorBench 3.2.0 at 73.4% against 70.5%, 70.0%, and 67.2%, zero scores on safeguard-intervened tasks with substitution by Claude Opus 4.8 and Claude Opus 5, 60% fewer cyber false positives and roughly 60% fewer Claude Code interventions per session and 85% less frequent biology safeguard firing, permitted vulnerability identification with exploit development still disallowed and penetration testing and exploit generation and binary-based vulnerability scanning routed to Opus models, Mythos 5.1's strongest-yet cyber capabilities remaining in the lower risk category of the Frontier Compliance Framework, a nearly 50% hit rate across 12 targets against a typical 10 to 15% and 10-times binding affinities on EGFR, Nipah G, and 15-PGDH with the stalk-targeting entry footnote at roughly 1.4 nM, the Venus elevation map from NASA Magellan radar covering a third of the planet at 2 to 3 kilometer detail with heights up to 25% more accurate under a Creative Commons license, seven open-source deep learning models sped up by as much as 2.5 times with 30 to 60% estimated GPU cost savings on genome-wide analyses, the anti-distillation restriction on new API accounts with existing accounts unaffected for now, the EFS architecture and its 100-plus customer development and supported platforms and phased fall rollout, the EU AI Act Code of Practice with 190 other signatories and watermarking for models released after August 2, 2026 with a private preview detection API, the claude-fable-5-1 API identifier with availability on Amazon Web Services and Google Cloud and Microsoft Azure, and Mythos 5.1's US-organization-only availability through the CVP and LSVP.

ASAP — AGI Soon As Possible

AI & tech,
read in depth

Beyond the headlines — into the context and the structure

AGI Soon As Possible · asapai.co.kr

← All posts