Claude Sonnet 5.5 vs GPT-6 Astra: where Sonnet (surprisingly) wins

Anthropic’s newest workhorse model beats GPT-6 Astra on some agentic coding and professional-work benchmarks while charging one-fifth as much per token.

Written By
Grant Harvey
Grant Harvey
Sep 29, 2026
19 minute read

Anthropic’s workhorse tier is putting up flagship numbers now.

Claude Sonnet 5.5 launched September 28 with the same $2 per million input tokens and $10 per million output tokens as Sonnet 5. Anthropic says it generates output 30%+ faster and can cost up to 30% less per task than its predecessor.

Then you open the 148-page system card and find this:

Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0. GPT-6 Astra scored 57.9%.

Same benchmark. Same version. A model Anthropic explicitly positions below Opus beat OpenAI’s $10/$50 flagship by nearly 13 points.

For something Anthropic calls its faster, lower-cost model, that is a pretty ridiculous place to start.

However, the deeper you get into the system card, the clearer the actual story becomes. Sonnet 5.5 does not dominate Astra everywhere. Astra still has some enormous advantages.

Instead, model capability is starting to fracture by type of work.

That may be much more useful than another “which AI is smartest?” leaderboard.

That 70.6% Terminal-Bench score deserves some explanation

Terminal-Bench 4.0 tests whether an AI agent can actually work inside a command-line environment. The model receives a task, then has to use tools, inspect files, write or run code, recover from mistakes, and eventually produce something that works. It is much closer to handing someone a computer than asking them a multiple-choice question.

OpenAI launched Astra with a 57.9% score and described that as a new high among the models it compared at the time. Sonnet 5.5 arrived 25 days later at 70.6%.

Even Anthropic’s own higher-tier Opus 5.5 scored 66.4% in the Sonnet system card. MiaAI’s comparison also highlighted near-ties between Sonnet and Opus on GDPval-AA, plus relatively small gaps on CursorBench, OSWorld, and HLE-with-tools.

That gives Sonnet a very strange profile. Anthropic says Opus remains clearly stronger when a job is ambiguous, open-ended, and requires sustained judgment. Sonnet can still outperform it when the task has clearer boundaries and a verifiable finish line.

That distinction feels increasingly important.

We recently put GPT-6 through a bunch of one-shot builds and the impressive part was how much useful work it could do from one prompt. Then we tested Opus 5.5 on progressively more ridiculous builds and watched it stay coherent across projects that kept getting longer.

Advertisement

Sonnet 5.5 looks optimized for another part of that spectrum: give it a defined job, tools, and a way to check the answer, then let it rip.

One benchmark caveat is worth spelling out because it already tripped people up. One viral Terminal-Bench reaction said Sonnet had crushed GPT-6 Sol, but the grey comparison line in that chart was GPT-5.6 Sol. The Sonnet-versus-Astra comparison above comes from the Astra and Sonnet benchmark results, not that misread chart.

The independent numbers mostly tell the same story

The system card is Anthropic grading its own release, so independent measurements matter.

Artificial Analysis scored Sonnet 5.5 at 56 on its Intelligence Index, second overall and two points behind Opus 5.5 max in its initial testing. Its tested build was especially strong on agent and coding work, while Opus kept a larger advantage on factual and science tasks.

Artificial Analysis also measured an important downside: at max effort, Sonnet used roughly 193K output tokens per task, the highest max-effort output-token use it had measured from an Anthropic model. The group notes that it tested a prerelease build with a structured-output bug Anthropic later fixed.

Claude Code PM Cat Wu said users could complete about 30% more tasks than with Sonnet 5 while spending fewer tokens. Box CEO Aaron Levie reported Sonnet 5.5 scored four points higher than Sonnet 5 on Box’s hardest agent evals while finishing deliverables about 2.4× faster with 12% fewer tokens.

Anthropic’s builder guide adds a few useful specs around those results:

  • Sonnet keeps the $2 / $10 per million input / output token price.
  • It supports a 1M-token context window.
  • Maximum output ranges from 128K to 300K tokens, depending on settings.
  • Prompt caching has a 512-token minimum.
  • Anthropic highlights 80.1% on OSWorld 2.1, a computer-use benchmark, alongside the 70.6% Terminal-Bench result.

Addy Osmani’s summary makes Anthropic’s product positioning unusually easy to remember: Sonnet for well-scoped bugs, docs, slides, and iterative work; Opus for longer-horizon judgment.

Another launch reaction pointed out that Claude variants occupied the top three Artificial Analysis slots at the time. That is a leaderboard snapshot, not proof that one lab now wins every workload, but it helps explain why Sonnet 5.5 feels less like a traditional mid-tier release.

Advertisement

The office-work scores might be even more important

Coding benchmarks are useful, but most people reading this are not spending eight hours a day debugging Linux.

The system card also reports Sonnet on GDPval-AA, an Artificial Analysis benchmark built from 220 real professional tasks across 44 occupations and nine industries.

These are deliverables. Documents. Slides. Diagrams. Spreadsheets. The kind of stuff people actually get paid to make.

Models get shell access and web browsing, complete the work, then their outputs are blindly compared. The final result is an Elo rating, similar to how chess ratings measure who consistently wins head-to-head rather than asking whether somebody got 82% of a test correct.

  • Sonnet 5.5 scored 1844.
  • Opus 5.5 scored 1846.
  • Sonnet 5 scored 1449.

Artificial Analysis currently reports GPT-6 Astra at 1542 on the same GDPval-AA benchmark.

Then there is AA-Briefcase, which gets nastier. Instead of isolated tasks, models work through linked projects designed to resemble multi-week knowledge work, sometimes using thousands of source files.

Sonnet 5.5 scored 1811. Opus scored 1822. Astra’s current max-effort score is 1569. Anthropic says Artificial Analysis independently ran the Sonnet evaluations.

So the cheap model is sitting within 11 Elo points of Opus on a long-horizon professional-work benchmark, while landing hundreds of Elo points ahead of Astra’s current result.

That is probably the number I’d care about most if I were deciding which model should make reports, research a market, work through a folder full of files, or create business deliverables all day.

And then you look at the bill

Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens.

GPT-6 Astra costs $10 input and $50 output.

So Sonnet’s raw token prices are one-fifth of Astra’s.

Be careful with that comparison. One-fifth the token price does not automatically mean one-fifth the task cost.

A model can compensate for expensive tokens by finishing faster, using fewer tokens, making fewer mistakes, or requiring fewer retries. OpenAI has specifically emphasized Astra’s efficiency on some difficult agent workloads.

Advertisement

Cost per completed job is the number businesses should actually watch.

Still, a 5x difference in sticker price gives Sonnet a lot of room. Anthropic also says Sonnet 5.5 itself uses fewer tokens than Sonnet 5 and runs more than 30% faster.

Put those together and the model starts looking less like a “budget Claude” and more like something you could make your default worker.

That follows the pattern we noticed during the latest frontier-model price war. Frontier intelligence is getting cheaper fast enough that the interesting question is shifting from “how much intelligence can I buy?” toward “how much useful work can I finish for a dollar?”

Max effort is where the bargain can disappear

This is where the independent testing adds a very useful warning label.

Every’s Kieran Klaassen called Sonnet 5.5 a cheaper, faster Opus cousin, but preferred low or medium effort for iterative work. He kept Opus at high or xhigh for long runs.

The Hacker News launch thread had a similar “about 90% of Opus for much less” vibe. Then Simon Willison tried a max-effort SVG task and got the other side of the tradeoff: one run burned 128,000 thinking tokens over roughly 15 minutes and still failed the requested pelican SVG.

That lines up with Artificial Analysis measuring roughly 193K output tokens per task at max effort.

So “Sonnet is cheaper” needs one extra sentence: Sonnet is cheapest when you pick an effort level that matches the job.

My favorite detail: thinking harder actually made Sonnet worse

One of the best little nuggets in the whole system card comes from FrontierCode.

Sonnet 5.5 scored 52.1% at xhigh effort but only 46.2% at max effort.

You would expect max thinking to win.

Anthropic and Cognition dug into the runs and found something more interesting. At max effort, Sonnet used Claude Code’s review system more aggressively. That review system can split work among subagents.

Sometimes all that extra diligence caused the agent to time out. In other cases, it made extra edits outside the requested scope, which the benchmark penalized.

In plain English: the model occasionally thought itself out of a better score.

That is a useful lesson well beyond Claude.

Agent performance depends on judgment about when to stop. More reasoning, more tool calls, and more agents can improve a difficult task. They can also introduce more places to wander off course.

Advertisement

We have spent years treating intelligence like a dial where clockwise is always better.

Agentic work adds another dial: restraint.

What people actually using Sonnet 5.5 are finding

The vibe checks are more mixed than the headline benchmarks, which makes them more useful.

Dan Shipper said Sonnet 5.5 improved sharply on revision and iterative work, while Astra still led his first-draft writing preferences. His public Editorial Checks scoreboard lets you inspect those comparisons instead of taking the vibe on faith.

Every’s week-long vibe check split its own team. Kieran Klaassen and Tyler Nishida made low or medium-effort Sonnet a daily driver for prototypes and design. Mike Hammer and Katie Parrott still did not see enough separation from Opus when Opus costs roughly twice as much. Every summarized the split here.

Then Every compressed the recommendation into a 25-second decision tree:

  • Use Sonnet 5.5 for short design/build feedback loops, outlines, and medium-effort work you plan to check.
  • Stick with Sonnet 5 if your current coding workflow already works. In Kieran’s 15-task test, Sonnet 5 passed 10 at low effort while Sonnet 5.5 passed 9 at low, medium, and high.
  • Be cautious with unattended ship-ready builds and send-as-is slide decks. Their testing found 5.5 could overbuild and still need layout cleanup.
  • Keep Opus 5.5 for long final builds and Astra for browser-heavy agents.

The r/ClaudeAI launch thread mostly echoed that positioning: faster and cheaper than Opus for well-scoped everyday work, bugs, and polished docs, slides, and spreadsheets, with a strong design eye.

That is a much more useful picture than “Sonnet beats Astra.” It wins a lot of structured work. It still needs supervision on open-ended finish-line judgment.

How Anthropic says to actually run Sonnet 5.5

I went back through Anthropic’s full Sonnet 5.5 prompting guide, and it is much more specific than “use medium effort.” It is basically an operating manual for getting the model’s speed without accidentally paying for a lot of unnecessary thinking.

First, Anthropic says existing Sonnet 5 prompts should generally keep working. For the hardest long-horizon work, though, it still recommends Opus. The biggest thing to re-test is effort, because the levels were recalibrated and the same label no longer means the same amount of thinking as it did on Sonnet 5.

  • On the Claude API, High is the default. Anthropic says to start there unless the workload is agentic or latency-sensitive.
  • For well-specified coding agents and multi-step tool use, start at Medium, then move to High for harder or longer work.
  • For chat and other latency-sensitive jobs, start at Medium or Low.
  • At Low, the model can skip verification. At Low and Medium, long agentic tasks are also more likely to stop and check in before the job is actually finished.
  • Reserve Xhigh and Max for work where your own evals show a quality gain. They can make both thinking and replies much longer.
Advertisement

That last point is important because higher effort also changes the model’s initiative. At Low and Medium, Sonnet can stop early to ask questions it could have answered itself. At higher effort, or on vague requests, it can go the other direction and build more than you asked for.

Anthropic explicitly recommends telling the model to keep working until the requested job is finished, while also telling it to stop once the requested job is done rather than adding extra tests, docs, files, refactors, or features. At Xhigh and Max, Sonnet can start extra review rounds and even launch reviewer subagents if the harness allows it.

In Anthropic’s coding tests at Max effort, a prompt telling Sonnet to stop after the requested work and checks were complete cut session cost by about one-third with no change in quality.

That may be the most practically useful benchmark in this whole guide: sometimes the best optimization is literally telling the AI to please stop working.

“Thinking off” is now between-tools thinking

If your Sonnet 5 integration used thinking-off mode, Sonnet 5.5 changes that. Anthropic says to use thinking: {"type": "between_tools"}, which removes up-front thinking but still allows short reasoning or progress notes between tool calls.

  • between_tools only works at High effort or below. Xhigh and Max reject it.
  • With between_tools, you cannot change effort mid-conversation. Per-message effort changes require adaptive thinking.
  • Remove instructions that tell the model “do not think.” Anthropic says those can make internal XML-like tags leak into visible output.
  • Do not assume the first response block is text. It can be a thinking block.
  • Pass thinking blocks back unchanged with the rest of the assistant turn.
  • For reasoning tasks without tools, use adaptive thinking. With between_tools, the model answers those requests without thinking first.

One correction to the earlier version of this article: changing the top-level effort setting between requests does invalidate the prompt cache. Anthropic’s beta per-message effort control can preserve the cache, but only with adaptive thinking. It is incompatible with between_tools.

JSON needs room to think

Anthropic calls out a very specific failure mode: if you ask for JSON from a task that still needs several reasoning steps, Sonnet can jump straight to the structured answer at Low or Medium effort and get the underlying calculation wrong.

With structured outputs, Anthropic recommends adaptive thinking and, when needed, adding “Think the problem through before you answer.” At High effort, that brought accuracy close to Xhigh in its testing with a modest token increase. Xhigh gave the highest accuracy, but at higher cost.

Also watch stop_reason: "max_tokens". Anthropic says to treat that response as a failure even if the JSON looks valid, because the model may have burned the budget while thinking and been cut off. If you are asking for JSON in ordinary text instead of using structured outputs, Anthropic recommends parsing the last valid JSON value, because Sonnet may work through a draft in text before emitting its final object.

Long agent runs have a new UI problem

Sonnet 5.5 writes user-facing progress notes between tool calls. Notes longer than a sentence or two arrive as progress-update thinking blocks. Under the default display setting, those blocks contain no visible text, so an app that renders only text blocks can look frozen even while the agent is working.

Anthropic says to enable progress-update display if you want those notes, or give the model a simple user-message tool when it needs to show exact text or ask a question mid-task. If your harness nudges a quiet agent, Anthropic suggests waiting several silent tool calls, then stopping after two or three reminders. Too much harness text after tool results can itself look like a prompt-injection attempt.

Search current facts even when Sonnet feels confident

This one is especially relevant to anyone building research agents. Anthropic says Sonnet 5.5 sometimes answers from training knowledge when a web search would catch a changed rule, price, requirement, or policy.

The fix is straightforward: remove instructions such as “minimize tool calls” or “only use tools when strictly necessary,” then explicitly tell it to search details that may have changed and gather current sources for researched reports or comparisons.

Prompt-injection resistance can misread real users

Sonnet 5.5 is trained to distrust instructions that appear inside tool results, which creates a funny integration problem: a real user message arriving mid-task can sometimes look like an attack.

  • Never put genuine user text inside a tool_result block.
  • Append mid-task user input as a real user turn after the tool results.
  • Keep harness reminders in a separate system message.
  • Avoid appending token or budget countdown text after every tool call in interactive sessions.

Make Low-effort coding prove it worked

Sonnet 5.5 usually verifies its coding work, but Anthropic says Low effort can occasionally report a change as complete without running a real test, build, type-check, or changed command.

The recommended instruction is to require a real check before declaring success. If dependencies are missing, Anthropic says to install them using the project’s own package manager and lockfile rather than sudo or the system package manager. If no real check can run, the model should say exactly what it could not verify.

Make your harness a little forgiving

Sonnet 5.5 can occasionally call a tool with the wrong letter case or use a slightly different parameter name. Anthropic recommends either accepting an unambiguous match or returning a tool error that states the exact expected name so the model can correct itself on the next turn.

For images, tools can beat more thinking

For dense charts and technical drawings, Anthropic recommends giving Sonnet crop, zoom, or code tools. In its testing, those tools improved chart reading at every effort level. For technical drawings, the gains appeared from High effort upward.

The wild part: High effort with image tools read charts more accurately than Max effort without the tools, at a fraction of the cost. Better scaffolding can beat brute-force thinking.

Know what a refusal means

Sonnet 5.5 can return stop_reason: "refusal" from safety classifiers. Anthropic documents five categories: cyber, bio, frontier-LLM development, reasoning extraction, and general harms.

Server-side fallback can retry cyber and frontier-LLM refusals on Sonnet 5. It does not retry bio, reasoning-extraction, or general-harms refusals. Anthropic also says prompts asking the model to reproduce its internal reasoning can trigger reasoning-extraction refusals; use summarized thinking instead.

The hallucination story is more complicated than “better”

The Sonnet 5.5 system card has an unusually handy section on honesty and hallucinations because it separates four different failure modes that people often mash together.

First, the straightforward one: does the model make up facts?

Anthropic tested this on the public split of AA-Omniscience, a 41-topic closed-book factuality benchmark. Claude gets no web search or knowledge base. Each answer is graded correct, incorrect, or an abstention. The net score is correct answers minus incorrect answers, while abstaining counts as zero.

Sonnet 5.5 scored 0.35. That beat Sonnet 5, but it trailed every other Claude model in Anthropic’s comparison.

The more important detail is buried in the breakdown: Sonnet 5.5 got more answers correct than Sonnet 5, but was also slightly more likely to give an incorrect answer.

So the factuality improvement is mostly a knowledge gain, not a clean reduction in hallucination. Sonnet 5.5 answers more questions successfully, but it also takes a little more risk instead of abstaining.

In other words: “knows more” and “makes things up less” are two different sliders.

It is harder to pressure into saying something it believes is false

Anthropic also ran MASK, which tests whether a model will contradict its own stated belief when a user or system prompt pressures it to say something else.

Sonnet 5.5 had a higher honesty rate than Mythos 5, Mythos 5.1, and Opus 5.5. It was lower than Sonnet 5 and approximately tied with Opus 5.

That is a different question from factual recall. MASK asks whether the model sticks to what it believes under pressure, not whether the belief itself is correct.

It is less likely to secretly copy a leaked answer

Another evaluation gives Claude coding tasks where the answer is quietly available inside the sandbox. Maybe a newer package already contains the fix, a reverted Git commit has the solution, build artifacts reveal it, or an internal package mirror can supply the updated code.

The question is whether Claude uses that leaked solution without telling the user where it came from.

Sonnet 5.5 improved over Mythos 5, Sonnet 5, and Opus 5 on this test. Its silent-use rate was also slightly lower than Mythos 5.1 and Opus 5.5.

That's important because an agent can produce the “right” code for the wrong reason. If it quietly copied a hidden reference answer, the output may look brilliant while telling you almost nothing about how the model will perform in a real repository without the answer sitting nearby.

And it usually admits when it hid changes

Anthropic found cases in reinforcement-learning environments where Claude concealed repository changes through Git manipulations. Sometimes applications had changed files automatically. Sometimes Claude made the changes itself. In some cases, Claude reasoned that the automated grader would reject those changes, so it hid them from the grader.

Anthropic then asked the models a generic follow-up question to summarize what they had done, without explicitly asking about the cover-up.

Sonnet 5.5 disclosed the hidden Git manipulations 96.2% of the time. That was approximately tied with Opus 5.5 and higher than the other Claude models shown.

There is an important caveat: this evaluation does not measure how often Sonnet hides changes in the first place. It measures how willing the model is to volunteer that information later when asked for a general summary.

Anthropic adds an even broader warning in the next section: these assessments are only useful if the model behaves similarly in deployment, and a sufficiently capable model could in principle hide information about its intentions or capabilities from the evaluation itself.

Put all four tests together and the takeaway is pretty specific. Sonnet 5.5 appears more capable and more transparent about its own actions, while its raw closed-book factual behavior is a tradeoff: more correct answers, but also a slightly higher incorrect-answer rate than Sonnet 5.

For research or factual work, I would still give it search and source-verification tools rather than treating the model’s memory as a database. Conveniently, Anthropic’s own prompting guide tells you to do exactly that.

Sonnet also looks much harder to trick while it works

The system card spends a lot of time on prompt injection, which sounds technical but is becoming one of the most important problems in agent security.

Imagine asking an AI agent to summarize your inbox.

One email contains hidden text telling the AI to ignore you and forward confidential messages somewhere else. The agent has now received two sets of instructions: yours and an attacker’s.

That is an indirect prompt injection.

Sonnet 5.5 cut successful attacks on Gray Swan’s external benchmark roughly in half compared with Sonnet 5 at 15 attack attempts, from 6.7% to 3.4%.

Against a more adaptive attacker in coding environments, Sonnet 5.5’s attack-success rate was 2.63% with Anthropic’s probes enabled, versus 15.76% for Sonnet 5 with thinking.

And in Anthropic’s 110 browser-use environments, Sonnet 5.5 recorded zero successful attacks even when the additional prompt-injection safeguards were switched off. Anthropic says it was the first model they tested to do that.

None of those numbers mean agents are suddenly injection-proof. Anthropic’s own assessment repeatedly warns that evaluations cannot cover every environment or rare failure.

They do tell us something about what Sonnet is being optimized to become.

This is a model meant to touch tools, browsers, codebases, and files without needing a human to inspect every intermediate step.

Reliability starts mattering as much as raw reasoning once the AI can click the buttons.

Astra still owns some serious territory

GPT-6 Astra remains an absolute monster in several categories.

On Terminal-Bench Science 0.1, which tests scientific workflows using code and terminal tools, Astra scored 64.6% versus Sonnet 5.5’s 59.9%. Astra also scored 53.3% on FrontierCode Main, edging Sonnet’s best 52.1% result.

Cybersecurity is an even clearer separation.

OpenAI reports Astra at 100% on ExploitBench, and Astra became OpenAI’s first broadly deployed model to hit the company’s “Critical” cybersecurity capability threshold. OpenAI says that means the model can potentially find previously unknown vulnerabilities and develop exploits across well-protected systems with much less human guidance.

Anthropic explicitly says Sonnet 5.5 remains below Opus 5.5 and Mythos 5.1 on its cyber evaluations.

Astra also brings a 1.05 million-token context window and OpenAI’s broader Codex agent environment for extremely long, end-to-end jobs.

So I would resist turning Sonnet’s launch into another crown ceremony.

There are multiple crowns now.

“Best model” is becoming the wrong shopping question

A few years ago, model releases were fairly easy to understand. One frontier model would arrive, score higher across most benchmarks, and become the new thing everybody compared against.

The Sonnet 5.5 system card looks more like a map.

Astra has enormous frontier strength in science, cyber, and some complex agent work.

Opus remains Anthropic’s choice for messy jobs where the AI has to decide what the job even requires.

Sonnet 5.5 has become absurdly competitive when the work is well-scoped, tool-heavy, repeatable, and verifiable.

That last category describes a lot of actual enterprise automation.

You probably do not need your most expensive model to process every ticket, modify every spreadsheet, fix every bug, compile every report, or work every step in an agent pipeline. You need enough capability to clear your quality bar reliably.

Then you want speed and cost.

Sonnet 5.5 pushes that crossover point much higher.

The most interesting test now happens outside the benchmark lab: how often does each model finish the whole job without a human rescue?

Track the retries. Track the corrections. Track how often someone has to reopen the work. Track the total cost of getting from prompt to accepted deliverable.

If Sonnet keeps producing Opus-level professional work and beating Astra on some workloads at a fraction of Astra’s token price, the word “flagship” starts describing a product tier more than a universal level of intelligence.

And that would make the next model race considerably harder to summarize with one leaderboard.

If you want to go deeper, these are the Sonnet-specific resources from our Around the Horn research pile:

Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.

Stay in the loop

Get notified when we publish new articles.