Skip to content
AlphaCorp AI
Wave of light particles flowing through faint circuit traces on a dark background
News22 min read

Claude Sonnet 5.5 Launch: Benchmarks, Pricing and Everything You Need to Know

Ignas Vaitukaitis, Founder & CEO of AlphaCorp AI

AI Agent Engineer ·

Claude Sonnet 5.5 Launch: Benchmarks, Pricing and Everything You Need to Know
On this page(32)
  1. Claude Sonnet 5.5 Launch at a Glance: Release Date, Positioning and Headline Numbers
  2. Claude Sonnet 5.5 Benchmarks Against Sonnet 5, Opus 5.5 and GPT-6 Sol
  3. How to read the FrontierCode row
  4. Where Sonnet 5.5 leads GPT-6 Sol
  5. Caveats attached to these scores
  6. Pricing Stays at $2/$10 While Cost per Task Falls
  7. Why the same price produces a smaller bill
  8. Speed and Effort Levels Determine What You Get per Dollar
  9. Why Max effort scored lower than Xhigh on FrontierCode
  10. Coding Performance Shows the Largest Generational Jump
  11. Knowledge Work, Computer Use and Design Close the Gap With Opus 5.5
  12. Design and document quality in practice
  13. Where Opus 5.5 remains stronger
  14. Technical Specifications and Platform Availability
  15. Where Sonnet 5.5 is available
  16. Breaking API Changes When Migrating From Sonnet 5
  17. Which changes cause the most trouble
  18. Recalibrate effort levels
  19. Safety Findings and the New Cyber, Biology and Distillation Safeguards
  20. What each safeguard does
  21. Refusal categories and verification programs
  22. Claude Sonnet 5.5 FAQ
  23. When was Claude Sonnet 5.5 released?
  24. How much does Claude Sonnet 5.5 cost?
  25. Is Claude Sonnet 5.5 better than Opus 5.5?
  26. Is Claude Sonnet 5.5 better than GPT-6 Sol?
  27. What is the context window of Claude Sonnet 5.5?
  28. Does Claude Sonnet 5.5 have a SWE-bench Verified score?
  29. Is Claude Sonnet 5.5 the default model in Claude Code?
  30. When will Claude Haiku 5.5 be released?
  31. Can I turn off thinking in Claude Sonnet 5.5?
  32. How to Choose and Start Using Sonnet 5.5 Today

Anthropic released Claude Sonnet 5.5 on September 28, 2026, at $2 per million input tokens and $10 per million output tokens, the same list price as Sonnet 5. It runs 30%+ faster and scores within a few points of Opus 5.5 on most published benchmarks. This breakdown of the Claude Sonnet 5.5 launch covers the scores, the real cost per task, the five breaking API changes and the new safeguards. All figures are current as of September 28, 2026.

Five numbers from Anthropic's September 2026 launch materials frame the release:

  • Agentic coding: Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 in 2026, against 10.3% for Sonnet 5 and 66.4% for Opus 5.5, per Anthropic's launch announcement.
  • Knowledge work: Sonnet 5.5 reaches 1,844 Elo on GDPval-AA v2.1 in 2026, two points behind Opus 5.5's 1,846, in results run by Artificial Analysis.
  • Cost per task: up to 30% lower than Sonnet 5 in Anthropic's 2026 testing, with the per-token price unchanged.
  • Capacity: a 1,000,000-token context window and 128,000 output tokens, per Anthropic's 2026 model documentation.
  • Migration: five breaking API changes from Sonnet 5, listed in Anthropic's developer notes dated September 28, 2026.

Every benchmark figure here is a vendor's own claim. Independent runs haven't confirmed them yet.

Claude Sonnet 5.5 Launch at a Glance: Release Date, Positioning and Headline Numbers

Anthropic released Claude Sonnet 5.5 on September 28, 2026, as the second model in the Claude 5.5 family. It runs more than 30% faster than Sonnet 5 and costs up to 30% less per task at the same list price. It replaces Claude Sonnet 5, which launched on June 30, 2026, and it arrived six days after Claude Opus 5.5 shipped on September 22, 2026.

The headline figures from Anthropic's Claude Sonnet 5.5 launch announcement are short enough to fit in one list:

  • Release date: September 28, 2026
  • List price: $2 per million input tokens and $10 per million output tokens, unchanged from Sonnet 5
  • Speed: output generated 30%+ faster than Sonnet 5 in Anthropic's 2026 testing
  • Cost per task: up to 30% lower than Sonnet 5 in the same testing
  • Agentic coding: 70.6% on Terminal-Bench 4.0, against 10.3% for Sonnet 5
  • Knowledge work: 1,844 on GDPval-AA v2.1, two points behind Opus 5.5's 1,846
  • Context window: 1,000,000 tokens, with up to 128,000 output tokens

Positioning matters as much as the numbers. Anthropic describes Sonnet 5.5 as a faster, lower-cost complement to Opus 5.5. Opus 5.5 is built for complex work that needs careful judgment. Sonnet 5.5 is aimed at well-scoped everyday tasks: fixing bugs, and producing documents, slides and spreadsheets that need little editing.

The family is still incomplete. Anthropic says Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will follow "in the coming weeks." Until then, Haiku 4.5 remains the cheapest and fastest model in the lineup.

One milestone is easy to skip past. Sonnet 5.5 is the first Sonnet model to beat Pokémon Red working only from screenshots, which Anthropic uses as a sign of long-horizon planning and image understanding.

Claude Sonnet 5.5 Benchmarks Against Sonnet 5, Opus 5.5 and GPT-6 Sol

Claude Sonnet 5.5 beats Sonnet 5 on all eight benchmarks Anthropic published in September 2026, lands within a few points of Opus 5.5 on most of them, and leads GPT-6 Sol on every benchmark where both have a score.

BenchmarkSonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Terminal-Bench 4.070.6%10.3%66.4%Not reported
FrontierCode 1.1 (Main)52.1% (Xhigh), 46.2% (Max)42.4%54.4%49.3%
CursorBench 4.055.5%34.1%57.8%Not reported
GDPval-AA v2.1 (Elo)1,8441,4491,8461,487
AA-Briefcase v1.1 (Elo)1,8111,3591,8221,483
Humanity's Last Exam (with tools)64.5%54.9%67.7%Not reported
OSWorld 2.1 (partial credit)80.1%57.0%81.8%Not reported
Chartography (no tools)61.6%15.6%64.4%53.6%
Grouped bar chart of five benchmark scores for Sonnet 5.5, Opus 5.5 and Sonnet 5, published September 2026. OSWorld 2.1 with partial credit: Sonnet 5.5 80.1%, Opus 5.5 81.8%, Sonnet 5 57.0%. Terminal-Bench 4.0: Sonnet 5.5 70.6%, Opus 5.5 66.4%, Sonnet 5 10.3%. Chartography without tools: Sonnet 5.5 61.6%, Opus 5.5 64.4%, Sonnet 5 15.6%. CursorBench 4.0: Sonnet 5.5 55.5%, Opus 5.5 57.8%, Sonnet 5 34.1%. FrontierCode 1.1 Main, with Sonnet 5.5 at Xhigh effort: Sonnet 5.5 52.1%, Opus 5.5 54.4%, Sonnet 5 42.4%.
Terminal-Bench 4.0 is the one benchmark where Sonnet 5.5 leads outright, at 70.6% against 66.4% for Opus 5.5 and 10.3% for Sonnet 5. Source: Anthropic, 2026.

How to read the FrontierCode row

Both FrontierCode figures belong to Sonnet 5.5. The model scores 52.1% at Xhigh effort, its best result, and 46.2% at Max effort. At Xhigh it beats GPT-6 Sol's 49.3% by 2.8 points and trails Opus 5.5's 54.4% by 2.3 points.

A higher effort setting producing a lower score looks like an error. It isn't one. FrontierCode checks whether a code change could be merged without human edits, and it penalizes out-of-scope changes even when they're helpful.

Where Sonnet 5.5 leads GPT-6 Sol

Four benchmarks have published scores for both models, and Sonnet 5.5 wins all four in the September 2026 results:

  • FrontierCode 1.1: 52.1% against 49.3%
  • GDPval-AA v2.1: 1,844 against 1,487
  • AA-Briefcase v1.1: 1,811 against 1,483
  • Chartography: 61.6% against 53.6%

The other four rows have no GPT-6 Sol figure. For Terminal-Bench 4.0, Anthropic's cost charts use GPT-5.6 Sol, because neither Terminal-Bench nor OpenAI had published a GPT-6 Sol result.

Caveats attached to these scores

Every number here comes from Anthropic's own launch materials, so treat the table as a vendor's claim until independent runs confirm it. Three footnotes deserve attention:

  1. Opus 5.5's 66.4% on Terminal-Bench 4.0 is its highest score, recorded at Xhigh effort. Sonnet 5.5 still beats it by 4.2 points.
  2. Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment of Sonnet 5.5 with a structured-outputs bug. Anthropic expects any effect to be small and to understate the model's performance. The bug is fixed.
  3. OpenAI fixed an image-understanding bug in GPT-6 Sol, and its GDPval-AA, AA-Briefcase and Chartography scores may predate that fix.

I'd read the third footnote as a reason to hold the GPT-6 Sol comparison loosely. The gaps on the two Elo benchmarks exceed 300 points, which a bug fix is unlikely to close. The 8-point Chartography gap is less settled.

Pricing Stays at $2/$10 While Cost per Task Falls

Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens as of September 28, 2026, the same list price as Sonnet 5. Anthropic's testing shows up to 30% lower cost per task. The per-token rate didn't move. The number of tokens needed to finish a job did.

Model (September 2026)Input per 1M tokensOutput per 1M tokens
Claude Haiku 4.5$1$5
Claude Sonnet 5.5$2$10
Claude Opus 5.5$4$20
Claude Fable 5.1$10$50
Grouped bar chart of Claude API list prices per 1 million tokens as of September 28, 2026, with input and output shown side by side. Claude Fable 5.1: $10 input, $50 output. Claude Opus 5.5: $4 input, $20 output. Claude Sonnet 5.5: $2 input, $10 output. Claude Haiku 4.5: $1 input, $5 output.
Sonnet 5.5 stays at $2 per million input tokens and $10 per million output tokens, half the $4 and $20 charged for Opus 5.5. Source: Anthropic, 2026.

Sonnet 5.5 sits at exactly half of Opus 5.5 on input and output, a ratio confirmed on Anthropic's published pricing page. The full rate card on the Sonnet 5.5 model overview in the Claude Platform docs adds the caching and batch rates:

  • Cache reads: $0.20 per million tokens
  • Cache writes with a 5-minute TTL: $2.50 per million tokens
  • Cache writes with a 1-hour TTL: $4 per million tokens
  • Batch API: 50% off input and output

Caching is where the two models converge. Cache reads cost $0.20 per million tokens on both Sonnet 5.5 and Opus 5.5, while cache writes cost $2.50 on Sonnet 5.5 against $5 on Opus 5.5.

Why the same price produces a smaller bill

Two things drive the savings, according to Anthropic. Sonnet 5.5 needs fewer tokens than Sonnet 5 to do the same work, and in head-to-head runs it batched tool calls together more often, which cut the number of steps per task.

That second point matters more than it sounds. In an agent loop, every extra tool call resends context and generates new output, and output is billed at five times the input rate. Fewer steps shrink the bill quickly.

"In our testing, it costs up to 30% less per task than its predecessor." (Anthropic, September 28, 2026)

Mind the words "up to." The 30% figure is a ceiling from Anthropic's own 2026 tests, and a short single-turn prompt has far fewer tool calls to batch than a long production AI agent run. Measure cost per completed task on your own workload before you rewrite a budget.

Speed and Effort Levels Determine What You Get per Dollar

Claude Sonnet 5.5 generates output 30%+ faster than Sonnet 5 in Anthropic's September 2026 testing, and the effort level you pick decides how much speed, quality and cost you get from it. Anthropic calls it the fastest Sonnet model to date.

Effort is a setting with five levels. Lower levels answer faster and use fewer tokens. Higher levels reason for longer and check their work more thoroughly.

Effort levelDefault inTypical use
LowNoneRoutine, high-volume work
MediumClaude Code and the Claude appsEveryday tasks
HighClaude Platform (API)Harder tasks that need more checking
XhighNoneBest FrontierCode result for Sonnet 5.5
MaxNoneLongest reasoning, highest cost per task

The split defaults catch teams out. A prototype built in Claude Code runs at Medium, and the same prompt sent through the API runs at High unless you set it yourself. Latency and cost per task will differ between the two, and the model hasn't changed.

Low effort is the bargain. On several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task, according to Anthropic's 2026 score-versus-cost charts. On Terminal-Bench 4.0, the Medium setting far exceeds Sonnet 5's best result of 10.3% for less than a tenth of the cost.

The relationship with Opus 5.5 changes as effort rises. Anthropic says Sonnet 5.5 complements Opus 5.5 best at lower settings, where it costs less per task. At higher settings, the two perform comparably at a similar cost. If you plan to run Sonnet 5.5 at Max all day, price the same job on Opus 5.5 first.

Why Max effort scored lower than Xhigh on FrontierCode

More effort doesn't always buy a better score. On FrontierCode 1.1 in September 2026, Sonnet 5.5 scored 52.1% at Xhigh and 46.2% at Max, a drop of 5.9 points. For reference, Sonnet 5 scored 42.4%, GPT-6 Sol 49.3% and Opus 5.5 54.4%.

Anthropic's explanation is specific. At Max, Sonnet 5.5 more often ran Claude Code's code-review skill, which splits a review across many subagents. In two cases Cognition examined, that led to a timeout or to edits beyond the task's scope, and FrontierCode penalizes both.

Bar chart of FrontierCode 1.1 (Main) scores from September 2026, ranked highest to lowest. Opus 5.5 54.4%. Sonnet 5.5 at Xhigh effort 52.1%, the highlighted bar. GPT-6 Sol 49.3%. Sonnet 5.5 at Max effort 46.2%. Sonnet 5 42.4%.
At Xhigh effort Sonnet 5.5 reaches 52.1%, its best FrontierCode result, while the longer-reasoning Max setting drops to 46.2%. Source: Anthropic, 2026.

Anyone who has run agents in production has seen this pattern. Extra diligence turns into scope creep.

Coding Performance Shows the Largest Generational Jump

Wireframe cubes of circuitry linked by glowing strands above a dark circuit-board floor
Built for production

What could a custom AI agent take off your plate?

We build production-grade AI systems that quietly handle the busywork, so your team can focus on the work that actually matters.

View Services

Coding is where Claude Sonnet 5.5 improves most over Sonnet 5, led by a Terminal-Bench 4.0 score that rose from 10.3% to 70.6% in Anthropic's September 2026 results. That's a gain of 60.3 points, close to seven times the earlier score. Terminal-Bench 4.0 measures how well a model completes multi-step professional tasks inside a command-line interface.

Three coding results stand out from the 2026 launch data:

  • Terminal-Bench 4.0: 70.6%, against 66.4% for Opus 5.5 at its best setting
  • FrontierCode 1.1 at High effort: 10 points above Sonnet 5 at the same setting, at about one fifteenth of the cost per task
  • CursorBench 4.0: 55.5%, against 34.1% for Sonnet 5 and 57.8% for Opus 5.5

Here's the caveat I'd attach to the Terminal-Bench number. A score of 10.3% was unusually weak for Sonnet 5, so part of the jump reflects a low starting point. A better yardstick is the previous flagship: Anthropic's Opus 5.5 launch page puts Opus 5 at 52.3% on the same benchmark in 2026. Sonnet 5.5 beats that by 18.3 points at a quarter of Opus 5's successor's price tier or less. Still a large gain. Just a less dramatic one.

Efficiency shows up beside the scores. Early testers told Anthropic that Sonnet 5.5 understands a codebase quickly, and in head-to-head runs it batched tool calls together more than Sonnet 5 did, which meant fewer steps per task.

"Claude Sonnet 5.5 cleared the same quality bar you'd expect from a higher-tier model, holding up on a system design audit and a data flow review." (Daniel Vogel, Chief Operating Officer, Epic Games, September 2026)

Other early-access customers gave Anthropic figures of their own in September 2026:

  1. SpaceXAI's Sualeh Asif cited the 55.5% CursorBench 4.0 score as "second only to Opus 5.5."
  2. Base44's Gabriel Grinberg reported that across 118 real app builds, Sonnet 5.5 produced apps scoring level with Opus 5, averaging 3.6 iterations per build.
  3. Unity's Sam Zhang reported a 90% completion rate on a multi-step Unity Editor and coding benchmark.
  4. Epic Games said the model managed tens of thousands of lines of gameplay system code and handled multi-hour tasks.

These are testimonials that Anthropic selected for its own launch page. Nobody has reproduced them independently, and each company tested on its own tasks. Treat them as a reason to run your own evaluation.

Knowledge Work, Computer Use and Design Close the Gap With Opus 5.5

Claude Sonnet 5.5 scores within 3 points of Opus 5.5 on knowledge work, computer use and chart reading in Anthropic's September 2026 results, at half the list price. On GDPval-AA v2.1, which tests real tasks across 44 occupations and nine major industries, the gap is 2 Elo points.

Benchmark (September 2026)Sonnet 5.5Opus 5.5Sonnet 5Gap to Opus 5.5
GDPval-AA v2.1 (Elo)1,8441,8461,4492
AA-Briefcase v1.1 (Elo)1,8111,8221,35911
OSWorld 2.1 (partial credit)80.1%81.8%57.0%1.7
Chartography (no tools)61.6%64.4%15.6%2.8
Grouped bar chart of Elo scores on two knowledge-work benchmarks from September 2026, comparing Sonnet 5.5, Opus 5.5 and Sonnet 5. GDPval-AA v2.1: Sonnet 5.5 1,844, Opus 5.5 1,846, Sonnet 5 1,449. AA-Briefcase v1.1: Sonnet 5.5 1,811, Opus 5.5 1,822, Sonnet 5 1,359.
On GDPval-AA v2.1 Sonnet 5.5 scores 1,844 against 1,846 for Opus 5.5, while Sonnet 5 sits 395 points lower at 1,449. Source: Anthropic, 2026.

The distance from Sonnet 5 is larger than the distance to Opus 5.5 on every row. Sonnet 5.5 sits 395 Elo points above its predecessor on GDPval-AA and 452 above it on AA-Briefcase. GPT-6 Sol scored 1,487 and 1,483 on those two tests in the same 2026 results.

Chartography is the startling one. Reading charts without tools went from 15.6% to 61.6% in a single generation. That fits with the Pokémon Red result, where Sonnet 5.5 became the first Sonnet model to finish the game working only from screenshots.

Design and document quality in practice

Testers also reported gains that benchmarks don't measure. They told Anthropic the model adds polish to user interfaces and follows slide templates closely enough that decks need minimal editing.

One internal test makes this concrete. Anthropic gave Sonnet 5.5 a public company's quarterly earnings materials, call transcripts and a slide template, then asked for a 10-slide operating review. Two experts judged the first draft ready to send as is.

Customer figures from September 2026 point the same way, with the usual caveat that Anthropic chose which quotes to publish:

  • Slack's Curtis Allen said Sonnet 5.5 beat Sonnet 5 on almost all offline Slackbot evals with about 14% fewer output tokens and no prompt changes.
  • Zendesk's Abhinay Kathuria said support tickets were processed 20% faster.
  • Atlassian's Jamil Valliani said Rovo Agents run up to 30% faster than on Sonnet 5.

Slack's detail about unchanged prompts matters to any team that has spent weeks on prompt engineering for an older model.

Where Opus 5.5 remains stronger

Near-parity on benchmarks doesn't make the two models interchangeable. Anthropic says that in its own testing and in external testing, Opus 5.5 remains clearly stronger at complex, open-ended work that requires sustained judgment. Its September 2026 release notes describe Sonnet 5.5 as a "faster, lower-cost complement" to Opus 5.5. Well-scoped tasks suit Sonnet 5.5. Ambiguous ones still favor Opus.

Technical Specifications and Platform Availability

Claude Sonnet 5.5 has a 1,000,000-token context window, a 128,000-token output limit and the model ID claude-sonnet-5-5. As of September 28, 2026, it is available on the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry. The figures below come from Anthropic's model documentation published that day.

SpecificationClaude Sonnet 5.5 (September 2026)
Model IDclaude-sonnet-5-5
Amazon Bedrock IDanthropic.claude-sonnet-5-5
Context window1,000,000 tokens
Max output128,000 tokens
Max output, Message Batches APIUp to 300,000 tokens with a beta header
Input and outputText and images in, text out
Knowledge cutoffJune 2026
ThinkingAdaptive, on by default
Comparative latencyFast
Earliest retirementSeptember 28, 2027

The 1M-token context window matches Opus 5.5, and so does the 128,000-token output limit. Buyers of the cheaper tier give up some judgment on hard problems, but they get the same capacity.

Anthropic's comparison table rates Sonnet 5.5 as "Fast" on latency. Opus 5.5 is rated "Moderate" and Haiku 4.5 "Fastest."

Where Sonnet 5.5 is available

Anthropic lists five access routes in its September 2026 documentation:

  • Claude API, open to all customers
  • Amazon Bedrock
  • Google Cloud
  • Microsoft Foundry
  • Claude Platform on AWS

Google Cloud and Microsoft Foundry use the same model ID as the Claude API. Bedrock adds the anthropic. prefix.

Two commitments matter for procurement. Anthropic won't retire the model before September 28, 2027, which gives teams at least a year of stable behavior to build against. Zero data retention is available as a deployment option, as it was for Sonnet 5 and Opus 5.5.

Breaking API Changes When Migrating From Sonnet 5

Migrating from Sonnet 5 to Claude Sonnet 5.5 involves five breaking API changes, so swapping the model ID alone can produce 400 errors. Anthropic documents them in its What's new in Claude Sonnet 5.5 developer notes, published September 28, 2026.

  1. Thinking can't be fully disabled. The "disabled" setting is gone. The lowest option is now between_tools, which still runs adaptive thinking but suppresses up-front reasoning before tool calls. If you run Sonnet with thinking off, switch to between_tools before you move.
  2. Forced tool use is removed. A tool_choice of "any" or a named "tool" returns a 400 error. Only "auto" and "none" work. Anthropic suggests strict tool use or structured outputs as replacements.
  3. Thinking blocks are bound to their origin. Sonnet 5.5 can read thinking blocks from Sonnet 5, Opus 4.8 and Haiku 4.5. It can't read blocks from Opus 5, Opus 5.5, or the Fable and Mythos lines, and no other model can read Sonnet 5.5's blocks.
  4. The old computer-use tool is rejected. The Claude API and Google Cloud refuse computer_20251124 and require computer_toolset_20260801. Bedrock still accepts the old tool.
  5. The advisor tool accepts fewer models. With Sonnet 5.5 as executor, the beta advisor tool no longer takes Opus 4.8, Opus 4.7 or Sonnet 5 as the advisor.

Which changes cause the most trouble

The second change will affect the most codebases. Forcing a named tool is a common way to get guaranteed JSON out of an extraction or classification pipeline. Every one of those calls fails on Sonnet 5.5 until you rewrite it.

The third change affects routing setups. If your system hands a conversation from Sonnet 5.5 to Opus 5.5 for a hard turn, the thinking blocks don't travel with it. The same applies to moving conversations between accounts, including switching accounts mid-session in Claude Code.

Setup you run todayAction before migrating
Thinking set to "disabled"Switch to between_tools
tool_choice of "any" or a named toolMove to strict tool use or structured outputs
Conversations passed between modelsCheck thinking-block compatibility
computer_20251124 on the API or Google CloudAdopt computer_toolset_20260801
Advisor set to Opus 4.8, Opus 4.7 or Sonnet 5Pick a supported advisor model

Recalibrate effort levels

Anthropic also advises against porting effort settings from Sonnet 5. Its documentation says effort levels "don't produce the same amount of thinking" on the new model. Rerun your evals at each level and choose again.

Safety Findings and the New Cyber, Biology and Distillation Safeguards

Claude Sonnet 5.5 matches or improves on Sonnet 5 on most alignment measures in Anthropic's September 2026 audit. It is also the first Sonnet model to launch with cyber safeguards, a fallback model and anti-distillation classifiers. Anthropic published a system card for the model on September 28, 2026.

The automated behavioral audit tests Claude across roughly 1,850 scenarios. Anthropic reports three findings from it:

  • Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse and honesty.
  • Sonnet 5.5 comes close to Opus 5.5 in how rarely it tries to escape its sandbox, and it is the least likely of any Anthropic model to probe the limits of its containers.
  • Opus 5.5 still performs slightly better across the full audit.

Anthropic found no evidence that Sonnet 5.5 pursues goals that conflict with the user's intention. It also says plainly that no set of evaluations reliably catches every failure. I'd keep both statements in mind.

What each safeguard does

SafeguardStatus in Sonnet 5.5Effect on users
CybersecurityNew for the Sonnet tierHigher-risk tasks visibly fall back to Sonnet 5
BiologySame as Sonnet 5Some microbiology and virology requests may be flagged in error
DistillationNew for the Sonnet tierClassifiers block reasoning extraction
WatermarkingIncludedInvisible text watermark, tied to EU AI Act compliance

The cyber safeguards exist because Anthropic rates the model's cybersecurity capability as comparable to Opus 5's. Routine bug finding and fixing still works. The fallback is visible, so a security team will know when a request was answered by Sonnet 5.

Distillation attacks use thousands of fake accounts to copy a model's capabilities into a version without safeguards. Sonnet 5.5 counters them with classifiers and with preserved thinking, which ties Claude's reasoning to the account that created it.

Refusal categories and verification programs

Refusals now arrive with a stop_details type. There are five categories, including "cyber", "bio", "frontier_llm" for requests that could help a competing lab build AI models, and "reasoning_extraction". Log them separately. A spike in one category tells you which safeguard your workload is hitting.

Two programs offer wider access. Organizations can apply to the Life Sciences Verification Program now. Anthropic says an expanded Cyber Verification Program for defenders will open soon, without giving a date.

One detail is absent from the launch announcement: the AI Safety Level assigned to Sonnet 5.5. Claude models since Opus 4 have run under the protections described in Anthropic's ASL-3 activation notice. Compliance teams should read the designation in the system card itself.

Claude Sonnet 5.5 FAQ

Claude Sonnet 5.5 launched on September 28, 2026, at $2 per million input tokens and $10 per million output tokens, with a 1,000,000-token context window. The short answers below cover the questions people search most often about the release.

When was Claude Sonnet 5.5 released?

Anthropic released Claude Sonnet 5.5 on September 28, 2026. It arrived six days after Claude Opus 5.5, which shipped on September 22, 2026, and about three months after Sonnet 5, which launched on June 30, 2026.

How much does Claude Sonnet 5.5 cost?

Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens as of September 28, 2026. Cache reads cost $0.20 per million tokens, and the Batch API takes 50% off input and output. The list price matches Sonnet 5, though Anthropic's 2026 testing shows up to 30% lower cost per task.

Is Claude Sonnet 5.5 better than Opus 5.5?

On most work, no. Sonnet 5.5 beats Opus 5.5 on one published benchmark, Terminal-Bench 4.0, with 70.6% against 66.4% in the September 2026 results. It trails Opus 5.5 on the other seven, usually by a few points. Anthropic says Opus 5.5 remains clearly stronger at complex, open-ended work that requires sustained judgment.

Is Claude Sonnet 5.5 better than GPT-6 Sol?

Sonnet 5.5 leads GPT-6 Sol on all four benchmarks where both models have a published score in September 2026: FrontierCode 1.1, GDPval-AA v2.1, AA-Briefcase v1.1 and Chartography. Those figures come from Anthropic's launch materials. Four other benchmarks have no GPT-6 Sol result, so the comparison covers half the table.

What is the context window of Claude Sonnet 5.5?

Claude Sonnet 5.5 has a context window of 1,000,000 tokens and a maximum output of 128,000 tokens. Both limits match Opus 5.5. The Message Batches API allows up to 300,000 output tokens with a beta header.

Does Claude Sonnet 5.5 have a SWE-bench Verified score?

Anthropic has published no SWE-bench Verified score for Claude Sonnet 5.5 as of September 28, 2026. Its launch benchmarks for agentic coding are Terminal-Bench 4.0, FrontierCode 1.1 and CursorBench 4.0. Any SWE-bench figure you see attached to Sonnet 5.5 didn't come from Anthropic's launch materials.

Is Claude Sonnet 5.5 the default model in Claude Code?

Anthropic hasn't announced Sonnet 5.5 as a default model in Claude Code or the Claude apps. Opus 5.5 became the Claude Code default with the September 22, 2026 client release. What is confirmed is the default effort level: Medium in Claude Code and the apps, High on the Claude Platform.

When will Claude Haiku 5.5 be released?

Anthropic says Claude Haiku 5.5 will arrive "in the coming weeks" and has given no date. It is built for high-volume and cost-sensitive applications. Until it ships, Haiku 4.5 at $1 per million input tokens and $5 per million output tokens stays the cheapest model in the lineup.

Can I turn off thinking in Claude Sonnet 5.5?

Thinking can't be fully disabled. The lowest setting is between_tools, which keeps up-front reasoning off before tool calls while adaptive thinking still runs. Teams that ran Sonnet 5 with thinking set to "disabled" need to change that setting before they migrate.

How to Choose and Start Using Sonnet 5.5 Today

Choose Sonnet 5.5 when the task is well scoped and cost per task matters, and plan the switch as a migration with testing. The Claude Sonnet 5.5 launch left a three-way choice for most teams:

  • Sonnet 5.5: bug fixes, documents, slides, spreadsheets and agent loops run at Low to High effort
  • Opus 5.5: ambiguous, open-ended work, or any job you'd otherwise run on Sonnet 5.5 at Max
  • Haiku 4.5: high-volume, cost-sensitive calls until Haiku 5.5 ships

Then take three steps, in order:

  1. Read Anthropic's migration guide and fix the five breaking changes, starting with forced tool use and disabled thinking.
  2. Change the model ID to claude-sonnet-5-5 in a staging environment.
  3. Rerun your evals at each effort level and record cost per completed task.
AlphaCorp AIonline
Let's talk

Curious what AI could do for your business?

No jargon and no hard sell. Just a friendly look at where AI fits, and where it doesn't.

View Services

Step three decides your budget. Anthropic's 30% saving is a ceiling from its own tests, and your workload will produce its own number.

If you'd like experienced builders to check your pipelines before the switch, AlphaCorp AI offers an AI integration audit that covers model migrations like this one.

Share
Newsletter · Weekly

Stay Ahead of AI

One email per week with the AI engineering insights, agent builds, and tools that actually matter.

No spamUnsubscribe anytimeFree forever

In every issue
  1. 01One agent build, taken apart step by step
  2. 02The tools that earned a place in our stack this week
  3. 03What broke in production, and what we changed

Written by Ignas Vaitukaitis, founder of AlphaCorp AI.

Wireframe cubes of circuitry linked by glowing strands above a dark circuit-board floor

Ready to Ship
Your AI System?

Book a free call and let's talk about what AI can do for your business. No sales pitch, just a real conversation.