What Is Claude Sonnet 5.5 in 2026?
Claude Sonnet 5.5 is the second model in Anthropic's Claude 5.5 family, announced on 28 September 2026. Anthropic describes it as a clear upgrade over Claude Sonnet 5 that runs more than 30 percent faster and costs up to 30 percent less for most work. It is available on all platforms including AWS, Google Cloud and Microsoft Azure, and on the Claude Platform as claude-sonnet-5-5. Claude Haiku 5.5 is said to follow in the coming weeks.
Anthropic positions it as the faster, lower-cost complement to Claude Opus 5.5. Opus 5.5 is for complex work requiring careful judgement; Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides and spreadsheets. Anthropic also says it has a sharp eye for design, which is an unusual claim to put in a model launch and a relevant one for anyone producing client-facing decks.
Why Does This Release Change Model Routing in 2026?
Because the mid tier has effectively caught the top tier on the work most businesses actually do, at half the input price. On Anthropic's own table, Sonnet 5.5 scores 1844 on GDPval-AA v2.1 against Opus 5.5's 1846, a gap of two points on a test of real-world work across 44 occupations. It also scores 70.6 percent on Terminal-Bench 4.0 against Opus 5.5's 66.4 percent. Sonnet 5.5 costs 2 and 10 US dollars per million tokens. Opus 5.5 costs 4 and 20.
For most of 2026 the routing advice has been simple: heavy tier for judgement, cheap tier for volume. This release complicates that in a useful way. A team whose main workloads are reporting, document production, bug fixing and well-scoped analysis now has a credible case for making Sonnet 5.5 the default and escalating to Opus 5.5 only for genuinely open-ended work. Anthropic is explicit that Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgement, and it says so directly rather than leaving the benchmark table to imply otherwise.
| Attribute | What Anthropic published in 2026 | Why it matters commercially |
|---|---|---|
| Announced | 28 September 2026, second model in the Claude 5.5 family | Six days after Opus 5.5. Haiku 5.5 still to come. |
| Price | 2 USD input, 10 USD output, 0.20 USD cache reads, 2.50 USD cache writes per 1M tokens | Identical to Sonnet 5. The saving comes from using fewer tokens, not a lower rate. |
| Cost per task | Up to 30 percent less than Sonnet 5 in Anthropic's testing | Token efficiency, not a rate cut, so your saving depends on your workload. |
| Speed | Generates output more than 30 percent faster than Sonnet 5 | Anthropic's fastest Sonnet model to date. |
| Positioning | Well-scoped everyday tasks, bug fixing, polished documents, slides and spreadsheets | This is most of a marketing or ops team's weekly output. |
| Effort defaults | Medium in Claude Code and the Claude apps, High on the Claude Platform | The default you inherit differs by surface, which changes your bill. |
| Safeguards | First Sonnet with Opus-class cyber safeguards; biology safeguards unchanged from Sonnet 5 | Higher-risk cyber tasks visibly fall back to Sonnet 5. |
| Migration note | If you run Sonnet with thinking off, switch to the new between_tools setting first | The one change that can break an existing integration. |
What Do the Claude Sonnet 5.5 Benchmarks Actually Say in 2026?
Anthropic published an eight row table against Sonnet 5, Opus 5.5 and GPT-6 Sol. The pattern is a very large jump over Sonnet 5 and near parity with Opus 5.5 on knowledge work, computer use and chart recognition, with Opus 5.5 still ahead on the harder coding evaluations. The table is reproduced as published.
| Benchmark (vendor-reported) | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 70.6% | 10.3% | 66.4% | not reported |
| FrontierCode 1.1 Main | 46.2% at Max, 52.1% at Xhigh | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | not reported |
| GDPval-AA v2.1 (knowledge work, Elo) | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 (knowledge work, Elo) | 1811 | 1359 | 1822 | 1483 |
| Humanity's Last Exam (with tools) | 64.5% | 54.9% | 67.7% | not reported |
| OSWorld 2.1 (computer use, partial) | 80.1% | 57.0% | 81.8% | not reported |
| Chartography (no tools) | 61.6% | 15.6% | 64.4% | 53.6% |
Two rows deserve a second look. The Terminal-Bench 4.0 jump from 10.3 to 70.6 percent is the largest single-generation move in any table we have covered in 2026, and it puts a mid-tier model above the flagship on that evaluation. The Chartography jump from 15.6 to 61.6 percent without tools matters to anyone who hands a model a dashboard screenshot and expects the numbers read back correctly.
The caveats belong with the numbers. Anthropic notes that benchmark scores capture only one facet of a model's capabilities, and that in its own testing and that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgement. Where GPT-6 Sol figures were not published for a benchmark, Anthropic reports GPT-5.6 Sol instead and says so. Every figure is Anthropic's own.
How Much Does Claude Sonnet 5.5 Cost in 2026?
Claude Sonnet 5.5 costs 2 US dollars per million input tokens and 10 US dollars per million output tokens, with cache reads at 0.20 and cache writes at 2.50. That is identical to Sonnet 5. The reduction Anthropic claims, up to 30 percent less per task, comes from the model needing far fewer tokens to do the same work rather than from any change to the rate card.
| Price per 1M tokens | Claude Sonnet 5.5 | Claude Opus 5.5 | Difference |
|---|---|---|---|
| Cache reads | 0.20 USD | 0.20 USD | Same |
| Cache writes | 2.50 USD | 5.00 USD | Sonnet 5.5 is half |
| Input tokens | 2.00 USD | 4.00 USD | Sonnet 5.5 is half |
| Output tokens | 10.00 USD | 20.00 USD | Sonnet 5.5 is half |
The cost-per-task claims are more interesting than the rate card. Anthropic reports that on several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task. On FrontierCode at High effort it scores 10 points higher than Sonnet 5 at the same setting for roughly a fifteenth of the cost per task. Both figures depend on effort settings and harness configuration, which is exactly why the number worth acting on is the one you measure yourself. Our model pricing comparison holds the full cross-vendor table.
Anthropic states the default effort is Medium in Claude Code and the Claude apps, and High on the Claude Platform. If your team works in the apps and your production integration runs on the Platform, the same task is being run at two different settings, with different token counts and different costs. Set effort deliberately per workflow rather than inheriting whatever the surface chose.
What Does Sonnet 5.5 Change for Knowledge Work in 2026?
It closes most of the gap to the flagship on the work a business actually pays people to do. On GDPval-AA v2.1, which tests real-world tasks across 44 occupations and nine major industries, Sonnet 5.5 scores about 400 points above Sonnet 5 and lands nearly level with Opus 5.5. On AA-Briefcase v1.1 the pattern repeats. Anthropic also says it clearly outperforms both Sonnet 5 and GPT-6 Sol on long-horizon knowledge work.
The most concrete example in the announcement is a document test rather than a benchmark. Anthropic gave the model a public company's quarterly earnings materials and call transcripts along with a slide template, and asked for a ten-slide operating review. Two experts judged the first draft ready to send as is. For any team that produces recurring client or board reporting from source documents, that is the claim to test first, because it describes the whole job rather than a step of it.
"Without changing any of our prompts, Claude Sonnet 5.5 did better than Sonnet 5 on almost all of our offline Slackbot evals, in fewer steps and with about 14% fewer output tokens."
Curtis Allen, Principal Engineer, Slack, as quoted by Anthropic
Anthropic also reports that early testers found it a more natural conversational partner, that it adds polish to user interfaces, and that it can follow slide templates to produce decks needing minimal editing. Template fidelity is the difference between a deck you send and a deck you rebuild.
How Good Is Claude Sonnet 5.5 at Coding in 2026?
Strong enough that the price gap to the flagship is now the deciding factor for a lot of routine engineering work. Anthropic reports that on CursorBench, which tests models on tasks from real Cursor coding sessions, Sonnet 5.5's best score lands within about two points of Opus 5.5. Early testers said it understood a codebase quickly and batched tool calls together more than Sonnet 5, producing fewer steps and lower costs.
"In Epic's early testing, Claude Sonnet 5.5 cleared the same quality bar you'd expect from a higher-tier model, holding up on a system design audit and a data flow review. The new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks, and delivered with less prescriptive prompting."
Daniel Vogel, Chief Operating Officer, Epic Games, as quoted by Anthropic
Anthropic also notes it is the first Sonnet model to beat Pokemon Red working only from screenshots, which sounds like a novelty and is actually a long-horizon vision and planning result. For marketing teams the practical read is on landing pages, tracking fixes and internal tooling: fewer steps means cheaper work and less to review.
What Do the Safety and Safeguard Changes Mean in Practice in 2026?
Sonnet 5.5 is the first Sonnet model to ship with Opus-class cyber safeguards, because Anthropic says its cyber capabilities are comparable to Opus 5's. Routine software development is unaffected, but higher-risk cybersecurity tasks visibly fall back to Sonnet 5. Biology safeguards are unchanged from Sonnet 5, and Anthropic notes that while they target harmful requests, some microbiology and virology requests may be flagged in error.
- Alignment: on an automated behavioral audit of roughly 1,850 scenarios, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse and honesty. Anthropic found no evidence that it pursues goals conflicting with the user's intention.
- Containment: on newer containment evaluations it comes close to Opus 5.5, and Anthropic says it is the least likely of any of its models to probe the limits of its containers. Opus 5.5 still performs slightly better across the full audit.
- Distillation: it is the first Sonnet model to launch with safety classifiers that prevent reasoning extraction, and it expands preserved thinking so Claude's thinking cannot be decoupled from the account that created it.
- Data: available with zero data retention, as with Opus 5.5 and Sonnet 5.
Anthropic repeats the caveat it used with Opus 5.5, and it is worth keeping: no set of evaluations reliably catches every failure, and Sonnet 5.5 may have tendencies it has not found. Human approval gates on consequential actions remain the right design.
The two changes that can break an integration
First, if you run Sonnet with thinking switched off, you need to move to the new between_tools setting, which keeps up-front thinking off, before switching to Sonnet 5.5. Anthropic publishes a migration guide for this. Second, preserved thinking now ties thinking to the account that created it, which affects anyone who moves conversations between accounts, including switching accounts mid-session in Claude Code. Anthropic says most developers will not notice, but if either pattern is in your stack, check it before you migrate.
How Should Marketing and Growth Teams Use Sonnet 5.5 in 2026?
Make it the default and escalate rather than the exception you reach for. The named strengths, well-scoped tasks, bug fixing, polished documents, slides and spreadsheets, plus design sense and template fidelity, map almost exactly onto a marketing team's recurring output. At half the price of the flagship and with near-parity on knowledge work, the burden of proof has shifted to justifying Opus 5.5 rather than justifying Sonnet.
| Workflow | Tier in 2026 | Human checkpoint |
|---|---|---|
| Recurring client or board reporting from source documents | Sonnet 5.5. This is the ten-slide operating review test. | Verify every figure against the source. |
| Decks against a house template | Sonnet 5.5. Template fidelity is a named strength. | Brand and claims review. |
| Landing pages, tracking fixes, internal tooling | Sonnet 5.5 at Low or Medium effort. | Code review and staged deploy. |
| Well-scoped analysis and summarisation at volume | Sonnet 5.5. Cheaper per task and faster. | Spot-check a sample. |
| Chart and dashboard reading | Sonnet 5.5. Chartography moved from 15.6 to 61.6 percent. | Check the numbers, not just the narrative. |
| Open-ended strategy, ambiguous briefs, final judgement | Opus 5.5. Anthropic says so explicitly. | Not an automation candidate. |
| Citation-critical research | Opus 5.5, which cleared Anthropic's citation-checked bar 16 times in 18. | Verify sources regardless. |
| Very high-volume classification and extraction | Wait for Haiku 5.5, or use a volume tier now. | Sampled audit. |
What Are the Common Mistakes to Avoid With Sonnet 5.5 in 2026?
- Expecting a cheaper rate card. The price is identical to Sonnet 5. The saving is token efficiency, so it depends entirely on your workload.
- Reading the Terminal-Bench result as "Sonnet beats Opus". On one evaluation it does. Anthropic states plainly that Opus 5.5 is stronger on complex open-ended work.
- Inheriting the effort default. Medium in the apps, High on the Platform. Same task, different bill.
- Migrating with thinking switched off. Move to
between_toolsfirst or the switch will not behave as expected. - Ignoring preserved thinking. If you move conversations between accounts or switch accounts mid-session in Claude Code, check the docs before migrating.
- Assuming cyber work carries over. Higher-risk cybersecurity tasks visibly fall back to Sonnet 5.
- Dropping human review because alignment improved. Anthropic itself says no evaluation set catches every failure.
Key Takeaways for 2026
Sonnet 5.5 is the release that makes the mid tier the sensible default for most business work, which is a bigger practical change than a new flagship usually delivers.
- Announced 28 September 2026 as the second Claude 5.5 model. Available on all platforms as
claude-sonnet-5-5. Haiku 5.5 is still to come. - Priced identically to Sonnet 5 at 2 and 10 US dollars per million tokens, with cache reads at 0.20 and cache writes at 2.50. Up to 30 percent cheaper per task through token efficiency.
- Output is more than 30 percent faster, making it Anthropic's fastest Sonnet to date.
- Scores 70.6 percent on Terminal-Bench 4.0 against Sonnet 5's 10.3 and Opus 5.5's 66.4, and 1844 on GDPval-AA against Opus 5.5's 1846.
- Anthropic states Opus 5.5 remains clearly stronger on complex, open-ended work requiring sustained judgement.
- First Sonnet with Opus-class cyber safeguards, with higher-risk cyber tasks visibly falling back to Sonnet 5, and first Sonnet with reasoning-extraction classifiers.
- Two migration traps: the new
between_toolssetting if you run thinking off, and preserved thinking if you move conversations between accounts.
Distk helps growth teams across India and internationally set the routing rules between tiers, find the cheapest effort level that still passes their own evaluation set, and place human checkpoints where a wrong number would actually cost something. If you are deciding what moves to Sonnet 5.5 and what stays on a flagship, that is the mapping we start with.
Sources
- Anthropic, Introducing Claude Sonnet 5.5, 28 September 2026. Every price, benchmark figure, safeguard detail and customer quote in this guide comes from that announcement and is vendor-reported unless attributed to a named third party.