AI Model Guide

Claude Sonnet 5.5 in 2026: The Mid Tier Just Caught the Flagship

Anthropic's second Claude 5.5 model costs exactly what Sonnet 5 cost, runs 30 percent faster, and lands two Elo points behind Opus 5.5 on real-world work at half the price. That changes which tier your workflows should default to.

Distk Editorial Sep 2026 13 min read

Claude Sonnet 5.5 is the second model in Anthropic's Claude 5.5 family, announced on 28 September 2026 and available on all platforms as claude-sonnet-5-5. Pricing is unchanged from Sonnet 5 at 2 US dollars input and 10 output per million tokens, with cache reads at 0.20 and cache writes at 2.50, but Anthropic reports up to 30 percent lower cost per task through token efficiency and more than 30 percent faster output. The headline numbers are a jump on Terminal-Bench 4.0 from Sonnet 5's 10.3 percent to 70.6 percent, which is above Opus 5.5's 66.4, and 1844 on GDPval-AA against Opus 5.5's 1846. Anthropic still states that Opus 5.5 is clearly stronger on complex open-ended work. Two changes can break an integration: the new between_tools setting if you run thinking off, and expanded preserved thinking if you move conversations between accounts. Haiku 5.5 is said to follow in the coming weeks.

What Is Claude Sonnet 5.5 in 2026?

Claude Sonnet 5.5 is the second model in Anthropic's Claude 5.5 family, announced on 28 September 2026. Anthropic describes it as a clear upgrade over Claude Sonnet 5 that runs more than 30 percent faster and costs up to 30 percent less for most work. It is available on all platforms including AWS, Google Cloud and Microsoft Azure, and on the Claude Platform as claude-sonnet-5-5. Claude Haiku 5.5 is said to follow in the coming weeks.

Anthropic positions it as the faster, lower-cost complement to Claude Opus 5.5. Opus 5.5 is for complex work requiring careful judgement; Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides and spreadsheets. Anthropic also says it has a sharp eye for design, which is an unusual claim to put in a model launch and a relevant one for anyone producing client-facing decks.

Why Does This Release Change Model Routing in 2026?

Because the mid tier has effectively caught the top tier on the work most businesses actually do, at half the input price. On Anthropic's own table, Sonnet 5.5 scores 1844 on GDPval-AA v2.1 against Opus 5.5's 1846, a gap of two points on a test of real-world work across 44 occupations. It also scores 70.6 percent on Terminal-Bench 4.0 against Opus 5.5's 66.4 percent. Sonnet 5.5 costs 2 and 10 US dollars per million tokens. Opus 5.5 costs 4 and 20.

For most of 2026 the routing advice has been simple: heavy tier for judgement, cheap tier for volume. This release complicates that in a useful way. A team whose main workloads are reporting, document production, bug fixing and well-scoped analysis now has a credible case for making Sonnet 5.5 the default and escalating to Opus 5.5 only for genuinely open-ended work. Anthropic is explicit that Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgement, and it says so directly rather than leaving the benchmark table to imply otherwise.

AttributeWhat Anthropic published in 2026Why it matters commercially
Announced28 September 2026, second model in the Claude 5.5 familySix days after Opus 5.5. Haiku 5.5 still to come.
Price2 USD input, 10 USD output, 0.20 USD cache reads, 2.50 USD cache writes per 1M tokensIdentical to Sonnet 5. The saving comes from using fewer tokens, not a lower rate.
Cost per taskUp to 30 percent less than Sonnet 5 in Anthropic's testingToken efficiency, not a rate cut, so your saving depends on your workload.
SpeedGenerates output more than 30 percent faster than Sonnet 5Anthropic's fastest Sonnet model to date.
PositioningWell-scoped everyday tasks, bug fixing, polished documents, slides and spreadsheetsThis is most of a marketing or ops team's weekly output.
Effort defaultsMedium in Claude Code and the Claude apps, High on the Claude PlatformThe default you inherit differs by surface, which changes your bill.
SafeguardsFirst Sonnet with Opus-class cyber safeguards; biology safeguards unchanged from Sonnet 5Higher-risk cyber tasks visibly fall back to Sonnet 5.
Migration noteIf you run Sonnet with thinking off, switch to the new between_tools setting firstThe one change that can break an existing integration.

What Do the Claude Sonnet 5.5 Benchmarks Actually Say in 2026?

Anthropic published an eight row table against Sonnet 5, Opus 5.5 and GPT-6 Sol. The pattern is a very large jump over Sonnet 5 and near parity with Opus 5.5 on knowledge work, computer use and chart recognition, with Opus 5.5 still ahead on the harder coding evaluations. The table is reproduced as published.

Benchmark (vendor-reported)Sonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Terminal-Bench 4.0 (agentic coding)70.6%10.3%66.4%not reported
FrontierCode 1.1 Main46.2% at Max, 52.1% at Xhigh42.4%54.4%49.3%
CursorBench 4.055.5%34.1%57.8%not reported
GDPval-AA v2.1 (knowledge work, Elo)1844144918461487
AA-Briefcase v1.1 (knowledge work, Elo)1811135918221483
Humanity's Last Exam (with tools)64.5%54.9%67.7%not reported
OSWorld 2.1 (computer use, partial)80.1%57.0%81.8%not reported
Chartography (no tools)61.6%15.6%64.4%53.6%

Two rows deserve a second look. The Terminal-Bench 4.0 jump from 10.3 to 70.6 percent is the largest single-generation move in any table we have covered in 2026, and it puts a mid-tier model above the flagship on that evaluation. The Chartography jump from 15.6 to 61.6 percent without tools matters to anyone who hands a model a dashboard screenshot and expects the numbers read back correctly.

The caveats belong with the numbers. Anthropic notes that benchmark scores capture only one facet of a model's capabilities, and that in its own testing and that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgement. Where GPT-6 Sol figures were not published for a benchmark, Anthropic reports GPT-5.6 Sol instead and says so. Every figure is Anthropic's own.

How Much Does Claude Sonnet 5.5 Cost in 2026?

Claude Sonnet 5.5 costs 2 US dollars per million input tokens and 10 US dollars per million output tokens, with cache reads at 0.20 and cache writes at 2.50. That is identical to Sonnet 5. The reduction Anthropic claims, up to 30 percent less per task, comes from the model needing far fewer tokens to do the same work rather than from any change to the rate card.

Price per 1M tokensClaude Sonnet 5.5Claude Opus 5.5Difference
Cache reads0.20 USD0.20 USDSame
Cache writes2.50 USD5.00 USDSonnet 5.5 is half
Input tokens2.00 USD4.00 USDSonnet 5.5 is half
Output tokens10.00 USD20.00 USDSonnet 5.5 is half

The cost-per-task claims are more interesting than the rate card. Anthropic reports that on several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task. On FrontierCode at High effort it scores 10 points higher than Sonnet 5 at the same setting for roughly a fifteenth of the cost per task. Both figures depend on effort settings and harness configuration, which is exactly why the number worth acting on is the one you measure yourself. Our model pricing comparison holds the full cross-vendor table.

The effort default that quietly changes your bill in 2026

Anthropic states the default effort is Medium in Claude Code and the Claude apps, and High on the Claude Platform. If your team works in the apps and your production integration runs on the Platform, the same task is being run at two different settings, with different token counts and different costs. Set effort deliberately per workflow rather than inheriting whatever the surface chose.

What Does Sonnet 5.5 Change for Knowledge Work in 2026?

It closes most of the gap to the flagship on the work a business actually pays people to do. On GDPval-AA v2.1, which tests real-world tasks across 44 occupations and nine major industries, Sonnet 5.5 scores about 400 points above Sonnet 5 and lands nearly level with Opus 5.5. On AA-Briefcase v1.1 the pattern repeats. Anthropic also says it clearly outperforms both Sonnet 5 and GPT-6 Sol on long-horizon knowledge work.

The most concrete example in the announcement is a document test rather than a benchmark. Anthropic gave the model a public company's quarterly earnings materials and call transcripts along with a slide template, and asked for a ten-slide operating review. Two experts judged the first draft ready to send as is. For any team that produces recurring client or board reporting from source documents, that is the claim to test first, because it describes the whole job rather than a step of it.

"Without changing any of our prompts, Claude Sonnet 5.5 did better than Sonnet 5 on almost all of our offline Slackbot evals, in fewer steps and with about 14% fewer output tokens."

Curtis Allen, Principal Engineer, Slack, as quoted by Anthropic

Anthropic also reports that early testers found it a more natural conversational partner, that it adds polish to user interfaces, and that it can follow slide templates to produce decks needing minimal editing. Template fidelity is the difference between a deck you send and a deck you rebuild.

How Good Is Claude Sonnet 5.5 at Coding in 2026?

Strong enough that the price gap to the flagship is now the deciding factor for a lot of routine engineering work. Anthropic reports that on CursorBench, which tests models on tasks from real Cursor coding sessions, Sonnet 5.5's best score lands within about two points of Opus 5.5. Early testers said it understood a codebase quickly and batched tool calls together more than Sonnet 5, producing fewer steps and lower costs.

"In Epic's early testing, Claude Sonnet 5.5 cleared the same quality bar you'd expect from a higher-tier model, holding up on a system design audit and a data flow review. The new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks, and delivered with less prescriptive prompting."

Daniel Vogel, Chief Operating Officer, Epic Games, as quoted by Anthropic

Anthropic also notes it is the first Sonnet model to beat Pokemon Red working only from screenshots, which sounds like a novelty and is actually a long-horizon vision and planning result. For marketing teams the practical read is on landing pages, tracking fixes and internal tooling: fewer steps means cheaper work and less to review.

What Do the Safety and Safeguard Changes Mean in Practice in 2026?

Sonnet 5.5 is the first Sonnet model to ship with Opus-class cyber safeguards, because Anthropic says its cyber capabilities are comparable to Opus 5's. Routine software development is unaffected, but higher-risk cybersecurity tasks visibly fall back to Sonnet 5. Biology safeguards are unchanged from Sonnet 5, and Anthropic notes that while they target harmful requests, some microbiology and virology requests may be flagged in error.

Anthropic repeats the caveat it used with Opus 5.5, and it is worth keeping: no set of evaluations reliably catches every failure, and Sonnet 5.5 may have tendencies it has not found. Human approval gates on consequential actions remain the right design.

The two changes that can break an integration

First, if you run Sonnet with thinking switched off, you need to move to the new between_tools setting, which keeps up-front thinking off, before switching to Sonnet 5.5. Anthropic publishes a migration guide for this. Second, preserved thinking now ties thinking to the account that created it, which affects anyone who moves conversations between accounts, including switching accounts mid-session in Claude Code. Anthropic says most developers will not notice, but if either pattern is in your stack, check it before you migrate.

How Should Marketing and Growth Teams Use Sonnet 5.5 in 2026?

Make it the default and escalate rather than the exception you reach for. The named strengths, well-scoped tasks, bug fixing, polished documents, slides and spreadsheets, plus design sense and template fidelity, map almost exactly onto a marketing team's recurring output. At half the price of the flagship and with near-parity on knowledge work, the burden of proof has shifted to justifying Opus 5.5 rather than justifying Sonnet.

WorkflowTier in 2026Human checkpoint
Recurring client or board reporting from source documentsSonnet 5.5. This is the ten-slide operating review test.Verify every figure against the source.
Decks against a house templateSonnet 5.5. Template fidelity is a named strength.Brand and claims review.
Landing pages, tracking fixes, internal toolingSonnet 5.5 at Low or Medium effort.Code review and staged deploy.
Well-scoped analysis and summarisation at volumeSonnet 5.5. Cheaper per task and faster.Spot-check a sample.
Chart and dashboard readingSonnet 5.5. Chartography moved from 15.6 to 61.6 percent.Check the numbers, not just the narrative.
Open-ended strategy, ambiguous briefs, final judgementOpus 5.5. Anthropic says so explicitly.Not an automation candidate.
Citation-critical researchOpus 5.5, which cleared Anthropic's citation-checked bar 16 times in 18.Verify sources regardless.
Very high-volume classification and extractionWait for Haiku 5.5, or use a volume tier now.Sampled audit.

What Are the Common Mistakes to Avoid With Sonnet 5.5 in 2026?

Key Takeaways for 2026

Sonnet 5.5 is the release that makes the mid tier the sensible default for most business work, which is a bigger practical change than a new flagship usually delivers.

Distk helps growth teams across India and internationally set the routing rules between tiers, find the cheapest effort level that still passes their own evaluation set, and place human checkpoints where a wrong number would actually cost something. If you are deciding what moves to Sonnet 5.5 and what stays on a flagship, that is the mapping we start with.

Sources

Claude Sonnet 5.5 in 2026: FAQs

What is Claude Sonnet 5.5 in 2026?

The second model in Anthropic's Claude 5.5 family, announced 28 September 2026. Anthropic calls it a clear upgrade over Sonnet 5 that runs more than 30 percent faster and costs up to 30 percent less for most work. It is the faster, lower-cost complement to Opus 5.5, available on AWS, Google Cloud, Azure and the Claude Platform as claude-sonnet-5-5.

How much does Claude Sonnet 5.5 cost in 2026?

2 US dollars per million input tokens and 10 per million output, with cache reads at 0.20 and cache writes at 2.50. That is identical to Sonnet 5. The up to 30 percent saving Anthropic reports comes from the model using fewer tokens per task, not from a lower rate.

Is Claude Sonnet 5.5 better than Opus 5.5?

On some benchmarks, yes. Sonnet 5.5 scores 70.6 percent on Terminal-Bench 4.0 against Opus 5.5's 66.4, and 1844 on GDPval-AA against 1846, which is near parity on real-world work at half the price. Anthropic states that Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgement.

What changed for developers moving to Sonnet 5.5?

Two things. If you run Sonnet with thinking switched off you must move to the new between_tools setting, which keeps up-front thinking off, before switching. And expanded preserved thinking means Claude's thinking cannot be decoupled from the account that created it, which affects moving conversations between accounts or switching accounts mid-session in Claude Code.

What are the Sonnet 5.5 safeguards in 2026?

It is the first Sonnet model with Opus-class cyber safeguards, because Anthropic says its cyber capabilities are comparable to Opus 5's. Routine development is unaffected but higher-risk cybersecurity tasks visibly fall back to Sonnet 5. Biology safeguards are unchanged from Sonnet 5. It is also the first Sonnet with classifiers preventing reasoning extraction.

Should marketing teams switch their default model to Sonnet 5.5?

For most recurring work, it is now the sensible default: reporting from source documents, decks against a template, well-scoped analysis, landing pages and chart reading. Escalate to Opus 5.5 for open-ended strategy, ambiguous briefs and citation-critical research. Test on your own workload before migrating production.

Default to the cheaper tier, escalate on purpose

Distk sets the routing rules between model tiers for your marketing workflows, finds the cheapest effort level that still passes your own evaluation set, and puts the review steps where a wrong number would actually cost you.

Start the conversation →