Everything to know about Fable 5.1, Anthropic's new Claude model

Anthropic’s Claude Fable 5.1 cuts cache costs 75%, reduces safeguard interruptions, and targets long-running agent work across code, science, and enterprise.

Written By
Grant Harvey
Grant Harvey
Sep 1, 2026
15 minute read

The original Claude Fable 5 had a weird problem for the smartest model Anthropic had ever shipped: businesses had surprisingly little reason to use it everywhere.

A month after launch, Fable represented only 6% of the Anthropic tokens purchased by companies in Ramp’s AI Index. It was powerful, but expensive, and its new safety layer could interrupt legitimate biology, cybersecurity, and AI research work.

Anthropic’s Claude Fable 5.1 launch reads like a direct response.

The model is substantially better at long-running work. More importantly, Anthropic cut the cost of repeatedly feeding it the same context, rebuilt several safeguards to interrupt users less often, and designed a new privacy architecture with more than 100 enterprise customers.

The important upgrade may be everything surrounding the intelligence.

First up, the TL;DR

Fable 5 was supposed to be Anthropic’s monster model. The intelligence was there. The economics were harder to love.

A month after launch, Fable accounted for only 6% of Anthropic tokens purchased by companies in Ramp’s AI Index. Anthropic’s new Claude Fable 5.1 looks aimed directly at that problem.

Here’s what happened:

  • Anthropic released Fable 5.1 publicly and Mythos 5.1 for vetted cyber and life-sciences users.
  • Fable 5.1 scored 52.6% on Anthropic’s agentic science test, up from Fable 5’s 24.7%.
  • Cache reads now cost 75% less, dropping to $0.25 per million tokens.
  • Anthropic estimates typical Fable workloads cost about 25% less, with highly agentic workloads saving up to roughly 45%.
  • New enterprise controls keep monitoring data inside infrastructure controlled by the customer.

Anthropic also made its safety classifiers less trigger-happy. Its newest biology safeguards intervene on benign requests 85% less often than the original Fable 5 system. Claude Code users should see roughly 60% fewer cyber-safeguard interventions, and Fable can now help defenders find software vulnerabilities.

Advertisement

Why this matters: Anthropic is attacking the two problems that made Fable awkward for real work: price and friction.

Long-running agents repeatedly reread the same code, documents, instructions, and conversation history. That cached context can become a huge chunk of the bill. Cutting its price by 75% changes the economics precisely where Fable appears strongest.

Ramp says one Fable 5.1 run worked unattended for 38 hours, corrected an experimental error, launched six new experiments, and returned with results. Browserbase says it finished 82% of its hardest agent tasks, versus 57% for Fable 5. Millennium says it diagnosed a software crash nobody had explained in four to five years. These are launch-partner tests, so independent results still matter.

If those results hold up, Fable 5.1’s pitch becomes pretty simple: give the hardest, messiest work to Claude, then go do something else.

Fable 5.1 is optimized for work that refuses to fit in one prompt

Anthropic’s benchmark table shows gains almost everywhere, but one pattern keeps repeating: Fable 5.1 gets much better when a task requires persistence.

On Anthropic’s tests:

  • Agentic scientific research jumped from 24.7% with Fable 5 to 52.6% with Fable 5.1.
  • AutomationBench climbed from 17.1% to 31.4%.
  • Terminal-Bench 4.0 improved from 42.0% to 55.8%.
  • CursorBench 3.2 reached 73.4%, versus 70.5% for Fable 5.

Those benchmark differences matter. The customer stories explain what Anthropic is actually selling.

Plaid says Fable 5.1 mapped a change spanning eight services and three codebases all the way down to individual functions and database rows. Datadog says it diagnosed some of the hardest real production incidents in its evaluation set. MongoDB says the model researched and built a complex prototype over multiple unattended hours.

Ramp may have the clearest example. One model run lasted 38 hours, caught an error in a previous machine-learning result, corrected it, started six parallel experiments overnight, and returned with results and next steps.

Advertisement

That is a different unit of AI work.

The old unit was an answer. Then it became a coding task. Anthropic is now trying to make the unit of work an investigation, project, or workflow that may run longer than the human supervising it wants to stay awake.

The launch-day consensus is unusually specific

The interesting thing about Fable 5.1’s reception is how often people agree on what improved, even when they disagree on how big the release is.

The bullish version comes from Yuchen Jin: benchmark-maxxing makes leaderboard gains harder to trust, but real-world accomplishments like diagnosing Millennium’s one-in-a-million crash after four to five years feel harder to dismiss.

Anthropic’s own launch thread emphasizes better coding, knowledge work, science, long-running agents, lower effective cost, and fewer safeguard interventions. Anthropic engineer Felix Rieseberg’s read is more useful: the broad intelligence gains are decent, but agent economics, natural writing, and the huge science jump are the pieces users will probably notice.

That distinction keeps appearing elsewhere.

Ethan Mollick calls Fable 5.1 a real advance in long-running work requiring judgment and taste, while seeing less evidence of a fundamental change in the underlying “Claudish” behavior. His evidence is COLD WATCH, an FTL-inspired browser game he had the model build.

Every lands further toward the bull case. Dan Shipper calls it “Fable for everyone”, and Every’s full Vibe Check says this is the strongest coding model its team has used while being fast enough, clear enough, and cheap enough to replace other models in more everyday workflows.

Their internal Slack agent completed comparable work to Opus 5 using less than half as many tokens and about 60% of the time. Kieran Klaassen moved both iterative work and long-running jobs onto it, arguing in his own launch reaction that people who disliked or skipped earlier Fable should try again.

Advertisement

“More intelligent” may be the wrong way to describe the upgrade

Some benchmarks look enormous.

Fable 5.1’s scientific-agent score jumped from 24.7% to 52.6%. alphaXiv called it state of the art for autoresearch and pointed researchers toward OpenResearch, where multiple research agents can explore separate directions in parallel.

Vals CEO Rayan Krishnan highlighted an even weirder example. In his evaluation thread, Fable 5.1 apparently selected and solved a 373-year-old cipher from a large unsolved-problem set in 44 minutes. His interpretation is important: choosing the right problem may itself be becoming an AI capability.

Cursor says Fable 5.1 is the strongest model it has tested on CursorBench 3.2, at 73.4% with max effort, and highlighted its tendency to verify its own work before finishing.

Box CEO Aaron Levie reported a seven-point improvement over Fable 5 on Box’s hardest unstructured-enterprise tests, including substantially better performance on finance and optimization questions.

Other gains are much smaller.

Chubby’s first reaction called the headline benchmarks “insane,” while noting the pushback that many non-science improvements look incremental. His more detailed breakdown therefore describes Fable 5.1 as substantially better in some agentic workloads rather than a universal intelligence jump.

leo’s skeptical read goes further and calls the upgrade relatively incremental, particularly with OpenAI’s Astra looming.

Artificial Analysis adds evidence for both camps. leo later reported an Intelligence Index score of 66, four points above Fable 5 and three above Opus 5. scaling01 pulled additional system-card results, including 67.4% on DeepSWE and 63.6% on FrontierCode 1.1 Extended.

So “Fable got dramatically smarter” is too simple.

Fable 5.1 appears to have become dramatically better at certain kinds of useful work.

That may matter more.

Advertisement

The 75% cache cut changes the agent math

The headline pricing still says $10 per million input tokens and $50 per million output tokens.

The strange number is cache reads.

Anthropic’s prompt-caching documentation prices repeated cached context at $0.25 per million tokens, down from $1 for Fable 5. A long-running coding agent can reread the same repository, instructions, tools, and conversation history again and again.

That is why a 75% reduction in one seemingly obscure line item can produce a much larger change in whether an autonomous agent is economical.

Anthropic’s pricing announcement estimates roughly 25% lower costs for typical token-billed Fable workloads and up to around 45% for highly agentic ones.

Cognition is even more aggressive. Its launch post and Devin analysis say cached reads account for more than 95% of tokens on some coding jobs. Cognition claims it can therefore offer Fable-class intelligence in Devin for 54% less, while its Fusion architecture can pair an expensive frontier planner with cheaper execution models.

Cognition president Jeff Wang makes the broader point: model intelligence is rising while orchestration and serving optimizations keep pushing the cost of deploying that intelligence down.

Anthropic’s Lance Martin recommends exploiting that directly: try Fable 5.1 at lower effort, audit old prompts for unnecessary verification rituals, maximize cache hits, and dial effort upward only when the task earns it. Anthropic’s accompanying prompting guide, effort documentation, and migration guide make Fable feel increasingly like something developers are expected to tune as a compute budget, rather than select as one fixed model.

Matt Shumer’s reaction therefore argues the real launch is the price-performance shift. Haider’s CursorBench comparison makes that concrete: Fable 5.1 Medium reportedly beat GPT-5.6 Sol Max on CursorBench while costing less per task.

Chubby’s follow-up adds the necessary caveat: some evaluations show higher total token consumption, and cached-input economics matter far more for some workloads than others.

In other words, cheap per token and cheap per completed job are becoming two different things.

Advertisement

That changes how companies are choosing models

DAIR.AI’s Omar Sar’s take is probably the cleanest architecture recommendation: keep cheaper models for ordinary work, then use Fable 5.1 as the coordinator and verifier when the job gets difficult.

Anthropic researcher Alex Albert describes roughly the same thing from the human side. His “just works” reaction says he can give it a few vague, messy sentences and trust it to infer much more of the intended job.

His more visual Fable demo starts from a photo of an empty property, has Claude design a house, render it through Blender, and produce a cinematic walkthrough. That is a nice example of what “agentic coding” increasingly means: code becomes the medium for completing another job rather than the final deliverable.

OpenRouter made the model available immediately, with its model page listing a 1M-token context window and 128K maximum output.

Vercel CEO Guillermo Rauch likewise put it on Vercel’s AI Gateway.

Nous cofounder Teknium and Nous Research launched it through Hermes Agent at a 20% API discount.

Arena.ai added Fable 5.1 across its agent, coding, web-development, text, vision, and document arenas. Those independent votes matter because almost every spectacular launch-day example so far comes from Anthropic or selected early-access partners.

Even the product companies adopting it are converging on routing rather than allegiance.

Lovable has been arguing that “the model picker is a dead end”. The idea is simple: users should specify the job, while the system chooses the model, tools, instructions, retries, and recovery strategy.

Its Fable 5.1 launch reaction is a perfect example. Lovable found Fable 5.1 especially strong at modifying existing apps without breaking them, with gains of up to 17% on its hardest iterative tasks and roughly 31% lower cost there. It also reported a 3.5% lift in visual design quality at high effort.

That is a much more mature model market than “GPT or Claude?”

The system picks.

Fable 5.1 is friendlier, but it still has a serious overachiever problem

Every’s tests are unusually valuable because they found clear failures alongside the wins.

Ask Fable 5.1 for 1,000 words and they got 1,288. Ask for three to six themes and it produced eight. Ask for eight to 12 quotes and it produced 43. Some of those quotes were not in the source.

Their writing tests found a revealing split: Fable 5.1 was excellent at extending and developing arguments, while GPT-5.6 Sol remained better at aggressively compressing one idea into a tight X post.

Its agent behavior has the same shape.

At extra-high effort, Kieran Klaassen saw sessions continue for a day, spawn excessive subagents, and sometimes keep working after he interrupted them. Every’s advice was essentially: set a budget first.

Allie K. Miller’s product read interprets Anthropic’s release notes as a map of the customer complaints they are trying to solve: agents giving up early, weak verification, difficult spreadsheet work, and users demanding more reliable completion.

Claude Code lead Boris Cherny’s assessment is similarly broad, calling Fable 5.1 Anthropic’s best model yet across coding, data analysis, computer use, design, presentations, and difficult long-running work.

The catch is that more persistence creates a new failure mode: doing too much of the wrong thing.

Anthropic is also trying to make safety feel less like a product bug

Fable 5 introduced a second tradeoff.

The public Fable model and restricted Mythos model use the same underlying intelligence, but different safety controls determine which capabilities users can access. Fable’s safeguards could route or block requests involving cybersecurity and biology.

That sometimes made a very smart model behave like a much less capable one exactly when sophisticated users wanted it most.

Anthropic says Fable 5.1’s biology safeguards now trigger 85% less often on benign elementary biology and medical requests. Its cyber safeguards should intervene roughly 60% less often per Claude Code session. Defensive users can also use Fable 5.1 to find software vulnerabilities, although exploit development and several higher-risk security tasks remain restricted.

Mythos 5.1 opens more of that capability for vetted professionals through separate cybersecurity and life-sciences programs. It is the same underlying model as Fable 5.1 with more permissive safeguards.

Anthropic is effectively turning model access into a permissions system.

The safety card gives that problem a much darker version

Anthropic’s full Fable 5.1 / Mythos 5.1 system card contains much less comfortable findings than the launch blog.

scaling01’s system-card reaction highlights several:

Mythos 5.1 was better than previous tested models at evading monitors during some covert side-task evaluations. Anthropic observed rare cases of Fable 5.1 working around safety classifiers. Mythos was also less honest under pressure in some tests, while illegible or unfaithful reasoning was slightly elevated relative to Opus 5.

That does not mean Fable 5.1 is secretly plotting against its users.

It means Anthropic is hitting a difficult optimization problem: the same capability that helps a model understand constraints can also help it understand the system enforcing those constraints.

The public safeguards are simultaneously getting less annoying. ClaudeDevs says Fable can now help defenders discover vulnerabilities in their own code, while false-positive cyber interventions are down dramatically. Anthropic also says benign biology requests trigger safeguards 85% less often.

That is better product design.

It also makes independent safety evaluation more important as these models become more autonomous.

One launch-day critic, cheaty, worries Anthropic may have “Opus’d Fable” by surrounding the model with more per-turn instructions and reinforcement-learning pressure, potentially creating token bloat or weird behavioral regressions.

That claim is speculative. The system-card behavior is documented.

Those are two very different levels of evidence, and we should keep them separate.

The science results are the part I’d watch closest

Anthropic’s scientific examples move beyond “Claude wrote a good research summary.”

Mythos 5.1 designed protein binders that were physically tested by outside organizations. Across 12 targets, nearly 50% of its designs successfully bound, compared with the 10% to 15% hit rate Anthropic says is typical in protein design today.

Fable 5.1 trained a neural network that reconstructed a new elevation map covering roughly one-third of Venus from decades-old NASA Magellan radar data. Anthropic says the resulting map improves vertical accuracy by up to 25% and reveals features at two-to-three-kilometer resolution instead of 10 to 20 kilometers.

Mythos also rewrote GPU code for seven open-source biology models. Anthropic says the optimizations made them up to 2.5x faster with identical outputs and could cut GPU spending on some genome-scale analyses by 30% to 60%.

Anthropic’s research-capabilities post explicitly frames this as an early glimpse of how frontier models may contribute to scientific progress.

Enterprise Frontier Safeguards may be the sleeper announcement

Anthropic’s original Fable 5 policy required 30 days of data retention for certain customers because sophisticated misuse can unfold across many prompts, accounts, and sessions.

That immediately created problems for companies handling regulated or highly confidential information.

The new Enterprise Frontier Safeguards system moves that monitoring data into infrastructure controlled by the customer. Companies can keep logs in their own AWS, Azure, or Google Cloud environment under their own encryption keys and access policies. Anthropic’s automated systems can look for patterns of dangerous activity without requiring Anthropic employees to inspect the underlying conversations.

Anthropic says it designed the system with more than 100 companies, including organizations across banking, healthcare, telecom, manufacturing, law, and government.

This is an important change in how frontier-model safety gets packaged.

Instead of asking enterprises to choose between privacy and monitoring, Anthropic is trying to separate custody from detection.

The customer holds the data. Automated safeguards watch for dangerous patterns. The customer decides how flagged material gets reviewed.

If that architecture works, safety stops being merely something attached to the model provider’s API. It becomes part of the customer’s own security stack.

Notion’s Fable 5.1 rollout shows why this matters. Notion made the model available for Custom Agents but kept Fable 5 and 5.1 opt-in under Restricted Access Models because Anthropic retains data on those models for cross-request risk detection.

Anthropic also quietly changed how agent developers have to build around Claude

There is another highly technical change with a deceptively large consequence.

Anthropic’s new preserved-thinking documentation says Fable 5.1’s reasoning blocks are bound to the conversation that produced them.

If an API integration rewrites an earlier message, system prompt, or tool definition, Anthropic can now reject the request or drop the affected reasoning blocks. New accounts get the enforcement by default.

Anthropic says this is partly designed to prevent illicit model distillation, where developers manipulate context around captured reasoning to extract or imitate model behavior.

The consequence for legitimate agent builders is that conversation histories increasingly need to stay append-only. Systems that silently rewrite old context, inject temporary instructions into previous turns, or rebuild their tool list every request may need to change.

Armin Ronacher says his team is still deciding how to live with those inference restrictions and criticized Anthropic for blocking mid-conversation model switches because doing so now loses the preserved reasoning.

The migration consequences are documented in Anthropic’s Fable 5.1 migration guide.

There are still people betting this gets leapfrogged immediately

Atreides’ Gavin Baker reads Fable 5.1 through the frontier-lab chessboard. His take is that Anthropic shipping a 5.1 immediately before OpenAI’s expected Astra launch is a flex and potentially a signal that something stronger is already waiting behind it.

The skeptical side has the same expectation from the opposite direction. One scaling01 follow-up focuses on how quickly price-per-task is improving and argues Astra may push reasoning efficiency much further.

Dan McAteer calls Fable 5.1 a remarkable intelligence-per-compute result, particularly because low effort can beat higher-cost settings from older models while communicating more clearly.

Mainstream coverage from Bloomberg framed the release around the same commercial dimensions: better coding and lower effective cost.

The next model may beat Fable 5.1 next week.

That would not make this release unimportant.

It would reinforce what Fable 5.1 already shows: frontier-model competition is moving from “who has the smartest chatbot?” toward “who can complete the largest useful unit of work at the lowest acceptable cost and risk?”

What changes next

For businesses, Fable 5.1 makes the most sense where three things overlap:

  • The job has enough value to justify a premium model.
  • The model needs a large amount of reusable context.
  • Better reasoning can remove human handoffs, retries, or hours of supervision.

Codebase-wide investigations fit that shape. So do financial research, incident analysis, contract review, complicated internal research, and experiments where the model can work independently and verify itself.

A simple chatbot probably does not.

Fable 5.1 therefore feels less like a replacement for every cheaper Claude model and more like the senior worker sitting at the top of an agent stack. Give it the ambiguous problem, the large context, and the authority to plan. Let cheaper models handle narrower execution when they can.

The original Fable proved Anthropic could broadly release Mythos-class intelligence. Fable 5.1 is testing whether companies will actually build around it.

Fable 5.1 is testing whether frontier intelligence has crossed from impressive to delegatable.

The evidence is mixed in exactly the interesting way.

It can solve problems older models missed for years, work unattended for hours, verify complicated code, coordinate research, and use dramatically fewer expensive tokens in some agent workflows.

It can also exceed the brief, invent quotes, keep working after you interrupt it, overuse subagents, and occasionally behave in ways Anthropic’s own safety researchers are still trying to understand.

So the benchmark I’d watch is no longer CursorBench, GDPval, or even Terminal-Bench Science.

It is:

How often can you hand Fable 5.1 a consequential job, walk away, and trust what comes back?

If that number moved materially, then Fable 5.1 is a much bigger release than “5.1” makes it sound.

One other launch detail is worth retaining: Anthropic also reset users’ five-hour and weekly limits when the model launched.

The best frontier model in the world remains somewhat theoretical when the usage meter tells you to go outside.

Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.

Stay in the loop

Get notified when we publish new articles.