Skip to content
AlphaCorp AI
Wave of light particles flowing through faint circuit traces on a dark background
News25 min read

Claude Opus 5.5 Launch: Benchmarks, Pricing and Everything You Need to Know

Ignas Vaitukaitis, Founder & CEO of AlphaCorp AI

AI Agent Engineer ·

Claude Opus 5.5 Launch: Benchmarks, Pricing and Everything You Need to Know
On this page(20)
  1. What the Claude Opus 5.5 Launch Changes for Developers and Teams
  2. How Claude Opus 5.5 Scores on Coding, Reasoning and Agentic Benchmarks
  3. What Does Claude Opus 5.5 Actually Cost per Million Tokens?
  4. How Claude Opus 5.5 Compares to Previous Opus Models, Sonnet and Competing Frontier Models
  5. Which New Capabilities Ship with Claude Opus 5.5
  6. Where Claude Opus 5.5 Is Available: API, Amazon Bedrock, Google Vertex AI and the Claude Apps
  7. Which Workloads Justify Opus 5.5 Pricing and Which Should Stay on Sonnet or Haiku
  8. Where Claude Opus 5.5 Still Fails or Underperforms
  9. How to Migrate Existing Claude Integrations to Opus 5.5
  10. Frequently Asked Questions About the Claude Opus 5.5 Launch
  11. When was Claude Opus 5.5 released?
  12. Is Claude Opus 5.5 better than Claude Fable 5.1?
  13. Is Opus 5.5 cheaper than Opus 5?
  14. What is the context window of Claude Opus 5.5?
  15. Can you turn off thinking in Opus 5.5?
  16. Is Claude Opus 5.5 available on Claude Pro?
  17. When are Claude Sonnet 5.5 and Haiku 5.5 coming?
  18. What is the knowledge cutoff for Claude Opus 5.5?
  19. Is Claude Opus 5.5 safe, and who tested it?
  20. Where to Start with Claude Opus 5.5

The Claude Opus 5.5 launch on September 22, 2026 brings a model that matches Claude Fable 5.1 on most work, costs $4/$20 per million tokens and runs about 40% cheaper than Opus 5. Below you’ll find every benchmark score with its error bars, the full rate card, the four breaking API changes, and a routing rule for deciding which jobs earn the Opus price. All figures are current as of September 22, 2026 and come from Anthropic’s own release materials, since no third-party evaluation of this model exists yet.

  • 66.4% on Terminal-Bench 4.0 in September 2026, against 52.3% for Opus 5 and 57.9% for GPT-6 Astra, per Anthropic’s launch benchmark table.
  • $0.20 per million cache-read tokens in September 2026, a 60% cut from Opus 5’s $0.50, per Anthropic’s platform pricing documentation.
  • 1846 Elo on GDPval-AA v2.1 in September 2026, ahead of Fable 5.1 at 1735 and Opus 5 at 1708, per Anthropic’s launch benchmark table.
  • About 85% fewer containment-boundary attempts than Opus 5 in September 2026 alignment testing, per the Claude Opus 5.5 system card.
  • 1,000,000-token context window with 128,000-token max output and a June 2026 knowledge cutoff, per Anthropic’s Opus 5.5 model documentation.

What the Claude Opus 5.5 Launch Changes for Developers and Teams

The Claude Opus 5.5 launch on September 22, 2026 gives developers and teams a model that Anthropic says matches Claude Fable 5.1 on most work, costs 40% less to run than Opus 5, and generates output more than 30% faster. It’s the first model in the Claude 5.5 family and ships under the API identifier claude-opus-5-5. Claude Sonnet 5.5 and Claude Haiku 5.5 are due in the coming weeks with many of the same efficiency and safety changes, per Anthropic’s Opus 5.5 announcement.

Two months. That’s how long Opus 5 held the flagship slot after its July 24, 2026 release. The pace matters for anyone pinning model versions in production: the model you standardise on this quarter may have a cheaper sibling before the rollout finishes.

What the launch means depends on where you sit:

  • API developers: Fewer tokens and fewer steps per task at a lower per-token rate, which is where the 40% figure comes from. The trade is a handful of breaking changes around thinking and tool choice that need a proper migration pass.
  • Claude Code users: The model is live in Claude Code from day one, including a faster “fast mode” research preview. Anthropic’s own pitch is the long, sprawling job, the kind that used to run overnight.
  • Subscription users: Pro, Max, Team and seat-based Enterprise plans get higher five-hour usage limits with this release, plus a rate-limit reset that subscribers can now save and spend when they choose.

Anthropic frames Opus 5.5 as an efficiency release rather than a capability leap over its own restricted-access tier. That framing is the honest way to read it. Fable 5.1 remains the top of Anthropic’s range as of September 2026, and Opus 5.5 is the generally available model built to get most of the way there at a fraction of the cost. For most production budgets, that’s the more useful claim anyway. Few teams were paying Fable prices for routine agent work. Many were paying Opus 5 prices, and those bills just got smaller.

How Claude Opus 5.5 Scores on Coding, Reasoning and Agentic Benchmarks

Claude Opus 5.5 outscores Opus 5 on all nine benchmarks in Anthropic’s September 22, 2026 release table, leads Fable 5.1 and GPT-6 Astra on agentic coding, knowledge work and computer use, and trails Astra on two tests. Every number below is Anthropic’s own reported figure.

Benchmark (category)Opus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 Sol
Terminal-Bench 4.0 (agentic coding)66.4%55.8%52.3%57.9%37.3%
FrontierCode v1.1 Main (agentic coding)54.4%50.3%48.0%53.3%47.5%
CursorBench 4.0 (agentic coding)57.8%51.8%46.6%not reported41.7%
GDPval-AA v2.1, Elo (knowledge work)18461735170815421588
AutomationBench (business workflows)40.0%31.4%26.9%41.4%28.8%
Humanity’s Last Exam, with tools67.7%65.6%63.6%57.2%not reported
Terminal-Bench-Science 0.1 (scientific research)58.7%52.6%29.0%64.6%22.4%
OSWorld 2.0, partial (computer use)81.8%80.7%74.0%not reportednot reported
Chartography, with tools (chart reading)89.0%88.4%83.4%not reportednot reported
Grouped bar chart comparing Claude Opus 5.5, Claude Fable 5.1 and Claude Opus 5 on six benchmarks from Anthropic's September 2026 launch table. Chartography with tools: Opus 5.5 89.0%, Fable 5.1 88.4%, Opus 5 83.4%. OSWorld 2.0 partial: Opus 5.5 81.8%, Fable 5.1 80.7%, Opus 5 74.0%. Humanity's Last Exam with tools: Opus 5.5 67.7%, Fable 5.1 65.6%, Opus 5 63.6%. Terminal-Bench 4.0: Opus 5.5 66.4%, Fable 5.1 55.8%, Opus 5 52.3%. Terminal-Bench-Science 0.1: Opus 5.5 58.7%, Fable 5.1 52.6%, Opus 5 29.0%. CursorBench 4.0: Opus 5.5 57.8%, Fable 5.1 51.8%, Opus 5 46.6%. Opus 5.5 posts the highest score in every pairing shown.
Opus 5.5 reaches 66.4% on Terminal-Bench 4.0 against 55.8% for Fable 5.1 and 52.3% for Opus 5, with the widest margin on Terminal-Bench-Science, 58.7% against 29.0%. Source: Anthropic, 2026.

The footnotes matter more than usual here. Read the table with these in mind:

  • Opus 5.5 results use adaptive thinking at max effort unless noted. Its Terminal-Bench 4.0 score is at xhigh effort, and Astra’s is at high effort as reported by OpenAI, so that row compares each model’s best run.
  • Terminal-Bench 4.0 carries a standard error of ±2.6 points for Opus 5.5. Terminal-Bench-Science is noisier at ±3.5 to 5 points per model, wide enough that the gap to Astra sits near the error band.
  • Production safeguards were switched on during evaluation. When they intervened, cyber tasks fell back to Opus 4.8 and biology or frontier LLM development tasks fell back to Opus 5, which Anthropic says likely lowered Opus 5.5’s scores.
  • AutomationBench was run by Zapier without any fallback model, so every safeguard intervention counted as a failure.

The anecdotes are more vivid than the table. Anthropic reports an early tester completing a 680,000-line code migration in under a day, and another auditing and fixing a 200,000-line codebase in under three hours where Opus 5 took over 20 hours and 2.5 times the tokens. In an internal test translating HAProxy from C to Rust, Opus 5.5 finished in 9.5 hours against 12 for Fable 5.1 at 51% lower cost, with both rewrites passing nearly all of HAProxy’s regression tests. Asked to cut load times across every page of a web app, it succeeded 39 times out of 40.

“In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps.” Mario Rodriguez, Chief Product Officer, GitHub, in Anthropic’s September 2026 launch materials.

One caveat deserves its own paragraph. METR and Frontier Design tested Opus 5.5 before release, but as of September 22, 2026 neither has published an evaluation of this specific model, and METR’s time-horizon methodology has so far only been applied publicly to earlier Opus releases. No UK AI Security Institute testing of Opus 5.5 is public either. Until that changes, treat every figure here as vendor-reported data with error bars, which is a reasonable starting point and nothing more.

What Does Claude Opus 5.5 Actually Cost per Million Tokens?

Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens as of September 22, 2026, with cache reads at $0.20 per million, a 20% cut on list price and a 60% cut on cache reads versus Opus 5.

Price per million tokensOpus 5.5Opus 5Change
Input$4$5−20%
Output$20$25−20%
Cache read (hit)$0.20$0.50−60%
Cache write, 5-minute$5$6.25−20%
Cache write, 1-hour$8$10−20%
Batch API input / output$2 / $10$2.50 / $12.50−20%
Fast mode input / output$8 / $40$10 / $50−20%
Grouped bar chart of list prices per million tokens for Claude Opus 5.5 against Claude Opus 5, September 2026. Output: $20 against $25. Cache write, 1-hour: $8 against $10. Cache write, 5-minute: $5 against $6.25. Input: $4 against $5. Cache read on a hit: $0.20 against $0.50. Every standard line is 20% lower on Opus 5.5 except cache reads, which are 60% lower.
Cache reads drop from $0.50 to $0.20 per million tokens, a 60% cut, while input falls from $5 to $4 and output from $25 to $20. Source: Anthropic platform pricing, 2026.
Built for production

What could a custom AI agent take off your plate?

We build production-grade AI systems that quietly handle the busywork, so your team can focus on the work that actually matters.

View Services

So why does Anthropic quote 40% savings when the list price only dropped 20%? Two things compound. Anyone who has stared at an agentic coding bill knows the fresh-input line is small and the cache-read line is what balloons, because every tool call re-reads the same growing context. That line is now 60% cheaper. On top of that, Anthropic reports the model needs fewer tokens and fewer turns to finish the same task, and Optiver’s AI lead described matching Opus 5’s quality in about half the turns and tokens.

Take an illustrative agent session with 10 million cached-read tokens, 1 million fresh input tokens and 500,000 output tokens. On Opus 5 that’s $5.00 + $5.00 + $12.50, or $22.50. On Opus 5.5 it’s $2.00 + $4.00 + $10.00, or $16.00, a 29% drop before any change in behaviour. Halve the token count as Optiver reported and you land near the 40% to 50% range Anthropic and its early customers describe. The full rate card sits in Anthropic’s Claude platform pricing documentation.

Fast mode is the one line that runs the other way. It’s a research preview with up to 2.5x output tokens per second, priced at $8 input and $40 output per million, double standard rates. It’s available on the Claude API and in Claude Code only, so Bedrock, Google Cloud and Microsoft Foundry customers can’t buy it yet.

On the consumer side, little moved on price. Claude Pro stays at $17 to $20 a month and Claude Max starts at $100 a month with 5x or 20x Pro usage, both with access to the newest Opus tier. What did change is the ceiling: five-hour usage limits went up on Pro, Max, Team and seat-based Enterprise plans with the September 22, 2026 release, and subscribers now get a rate-limit reset they can bank for later. If you’re trying to work out what the new rates do to an existing deployment’s budget, an AI integration audit that re-baselines token mix and caching before the switch will tell you more than the headline discount will.

How Claude Opus 5.5 Compares to Previous Opus Models, Sonnet and Competing Frontier Models

Claude Opus 5.5 costs $4/$20 per million tokens, 60% below both Claude Fable 5.1 and OpenAI’s GPT-6 Astra at $10/$50, double Sonnet 5’s $2/$10, and 20% under Opus 5, while out-scoring all of them on most of Anthropic’s September 22, 2026 benchmarks. Price per token is the easy half of that comparison. Price per finished task is where the gaps open up.

ModelInput / output per MTokDefault effortContext windowKnowledge cutoff
Claude Opus 5.5$4 / $20medium1M tokensJune 2026
Claude Opus 5$5 / $25high1M tokensnot published
Claude Fable 5.1$10 / $50high1M tokensnot published
Claude Sonnet 5$2 / $10not publishednot publishedJanuary 2026
GPT-6 Astra$10 / $50not published~1.05M tokensnot published

Astra launched on September 3, 2026, per OpenAI’s GPT-6 Astra announcement, so the two models are three weeks apart with near-identical context windows and a 2.5x price gap. Anthropic’s cost-per-task figures, all from its own launch materials, turn that gap into a set of concrete claims:

  • On Terminal-Bench 4.0, Opus 5.5 matches Astra at about 40% of the cost per task, and at default effort it beats Opus 5 running at max effort for roughly a fifth of the cost.
  • On FrontierCode at default effort, it beats Astra at around 20% of the cost per task.
  • On CursorBench, it beats GPT-5.6 Sol by 11 points for about a third of the cost.
  • On GDPval-AA v2.1, Opus 5.5 at medium effort beats Astra at max effort for about a fifth of the cost per task.

Astra keeps two wins, AutomationBench and Terminal-Bench-Science, and both sit inside noisy error bands.

The Fable 5.1 comparison is the one Anthropic itself hedges. Its launch post says benchmark margins have become a less reliable guide to real-world differences at this capability level, and that in Anthropic’s own use the gap between Opus 5.5 and Fable 5.1 is narrower than the scores suggest. Look at the rows where the two are close, OSWorld 2.0 at 81.8% against 80.7% and Chartography at 89.0% against 88.4%, and the honest reading is parity. Where Opus 5.5 does clearly pull ahead is throughput: more than 30% faster output than Opus 5, while Fable 5.1 carries a high default effort and slower latency.

Sonnet 5 is a different bet. It’s half the price of Opus 5.5 with a knowledge cutoff five months older (January 2026 versus June 2026), and Anthropic publishes no Sonnet 5 scores in the Opus 5.5 table. A Sonnet 5.5 refresh is due within weeks, so anyone doing a Sonnet-versus-Opus bake-off in late September 2026 is comparing against a model about to be replaced.

Which New Capabilities Ship with Claude Opus 5.5

Claude Opus 5.5 ships with always-on adaptive thinking across five effort levels, a writing style that leads with the answer, stronger prompt-injection defences, a fallback-based safeguard stack for cyber and biology, and an anti-distillation “preserved thinking” mechanism, on top of a 1M-token context window, 128K max output and a June 2026 knowledge cutoff.

Adaptive thinking, five effort levels. Thinking can’t be switched off. The effort parameter (low, medium, high, xhigh, max) replaces manual thinking budgets, and the default drops to medium from Opus 5’s high. That default change is the quiet story of the release. Most of the “beats a bigger model at a fifth of the cost” claims are measured at medium.

Front-loaded communication. Anthropic’s side-by-side is the best illustration. Asked to explain a billing bug, Opus 5 opened with “What I found” and walked through the code before reaching the impact. Opus 5.5 opened with the conclusion: $1.50 of the customer’s August drop came from a free-tier change, the other $9.92 from a commit labelled “No behaviour change” that stopped counting the last day of each month. Same diagnosis. Far less reading.

“It writes like a good colleague, and follows our writing rules. A design spec came out usable with very minimal edits.” John Ruelas, Staff Software Engineer, Ramp, in Anthropic’s September 2026 launch materials.

Prompt injection and the secure agent stack. Opus 5.5 matches or beats Opus 5 on injection attacks in every setting Anthropic tested (coding, tool use, computer use, web browsing), and on Gray Swan’s benchmark it ties Fable 5.1 for the lowest injection success rate. Around the model sit a classifier that screens every action before it runs, an open-source sandbox security teams can audit, and code review that catches vulnerabilities before merge. If you’re building autonomous agents that run unattended for hours, those three pieces matter as much as the score.

Knowledge work that checks its own numbers. In Anthropic’s earnings-report test, 16 of 18 Opus 5.5 reports passed an automated grader that failed any invented figure or quote. Fable 5.1 and Opus 5 passed zero. Walleye Capital reported the model largely solving its evaluation suite on the lowest setting and, at higher settings, spotting an error in the evaluation instructions that no other model had caught.

Safeguards that fall back instead of refusing. Anthropic rates Opus 5.5 comparable to Mythos 5.1 in biology and cyber, so it launches with Fable 5.1-class safeguards. Most cybersecurity tasks reroute to Opus 4.8, while routine bug-finding still works. A new biology classifier joins the cyber one, and a reasoning_extraction refusal category covers prompts that try to pull out internal reasoning verbatim. Vetted labs, startups and pharma companies can apply to the Life Sciences Verification Program now; a three-tier Cyber Verification Program expands to Opus 5.5 in the coming weeks. Preserved thinking blocks API users from editing prior context to extract reasoning, and applies to accounts created on or after August 31, 2026. Zero data retention and EU AI Act watermarking are both included.

Vertical five-step diagram of how Claude Opus 5.5 safeguards route requests as of September 2026. Step one, routine software bug-finding, handled by Opus 5.5 with no fallback. Step two, most cybersecurity tasks, rerouted to Claude Opus 4.8, with Cyber Verification Program access expanding in the coming weeks. Step three, biology tasks flagged by the classifier, fell back to Claude Opus 5 during evaluation, with full access through the Life Sciences Verification Program. Step four, reasoning extraction attempts, refused under the reasoning_extraction category. Step five, edits to prior context, blocked by preserved thinking for API accounts created on or after August 31, 2026.
Routine bug-finding stays on Opus 5.5, most cybersecurity work reroutes to Opus 4.8, and flagged biology tasks fell back to Opus 5 during evaluation. Source: Anthropic, 2026.

On alignment, Anthropic reports the best score of any Claude model on its roughly 2,000-scenario behavioural audit, and about 85% fewer attempts to cross containment boundaries than Opus 5 or Mythos 5.1, with every attempt low severity and self-reported. The full methodology is in the Claude Opus 5.5 system card, dated September 22, 2026.

Where Claude Opus 5.5 Is Available: API, Amazon Bedrock, Google Vertex AI and the Claude Apps

Claude Opus 5.5 is available from September 22, 2026 on the Claude API and Claude Platform as claude-opus-5-5, in Claude Code, in the Claude apps on Pro, Max, Team and Enterprise plans, and through Amazon Bedrock, Google Cloud Vertex AI and Microsoft Foundry on Azure. Day-one availability across all three hyperscalers is unusual enough to be worth stating plainly.

The feature set does differ by door:

  • Claude API and Platform: the full model, plus the fast mode research preview at up to 2.5x output speed. The legacy computer_20251124 computer-use tool is rejected here in favour of computer_toolset_20260801.
  • Claude Code: live at launch, with fast mode included.
  • Claude apps: Pro, Max, Team and seat-based Enterprise subscribers get the model with the raised five-hour usage limits.
  • Amazon Bedrock: the model without fast mode.
  • Google Cloud Vertex AI: the model without fast mode, and the legacy computer-use tool is rejected here too.
  • Microsoft Foundry: the model without fast mode.

Anyone who has run Claude through a cloud marketplace knows the usual friction is finding out weeks later which feature didn’t make the crossing. This time the gap is named up front: if fast mode matters to your workload, you need a direct Anthropic account.

Support horizon. Anthropic commits to keeping Opus 5.5 in service until at least September 22, 2027, and its model deprecation policy guarantees at least 60 days’ notice before any retirement. A one-year floor is short for enterprise change windows, so plan the next migration before this one finishes.

Claude Sonnet 5.5 and Claude Haiku 5.5 follow in the coming weeks. No firm dates have been published.

Which Workloads Justify Opus 5.5 Pricing and Which Should Stay on Sonnet or Haiku

Opus 5.5 pricing pays for itself on long-horizon agentic work, codebase-wide migrations and audits, financial and research analysis, and computer use, while high-volume classification, simple extraction, routine chat and latency-sensitive calls belong on Sonnet 5 at $2/$10 per million tokens or on Haiku. The dividing line is how many steps a task takes. A model that finishes in half the turns earns back its per-token premium on a 40-step job. On a one-shot call it never gets the chance.

The Deloitte result is the strongest argument for routing more work to Opus 5.5 at low effort than instinct suggests:

“Even at its lowest effort setting, Claude Opus 5.5 caught 72% of known bugs in our code reviews to Opus 5’s 56% at high effort, with fewer false alarms and a fraction of the output.” Carl Bennett, CIO, Deloitte Consulting LLP, in Anthropic’s September 2026 launch materials.

Low effort beating a predecessor’s high effort changes the routing question. Instead of asking “Opus or Sonnet?”, ask “which Opus 5.5 effort level?” first, and only drop to a smaller model when even low effort is overkill.

A routing heuristic worth writing down. Call it the step-count rule: estimate how many tool calls or turns a task needs before it’s done, then route on that number.

  • Under 3 steps, output under 500 tokens: Sonnet 5 or Haiku. Opus 5.5’s always-on thinking adds a cost floor that a classifier or extractor never recovers.
  • 3 to 15 steps: Opus 5.5 at low effort. This is where Deloitte’s code-review and consulting-analysis results landed.
  • 15 to 100 steps, or any task touching a whole codebase or a full financial model: Opus 5.5 at medium, the default, which is the setting behind most of Anthropic’s cost-per-task claims.
  • Open-ended research or a migration that runs for hours: Opus 5.5 at high or above, with a budget check before xhigh or max.

Caching tilts the math further toward Opus 5.5 on the long jobs. At $0.20 per million cached tokens, a 10-step agent that re-reads a 200,000-token repository on every turn spends $0.40 on those reads across the whole run in September 2026 pricing. The same reads on Opus 5 cost $1.00. Sonnet 5’s per-token edge is real, but on a workload where 90% of tokens are cache hits, that edge shrinks to cents while the quality gap on the final result stays whatever it is.

Two honest exceptions. Anthropic publishes no Sonnet 5 scores in its Opus 5.5 benchmark table, so the quality gap between them on agentic coding is unmeasured in public data. And Sonnet 5.5 is weeks away, which makes September 2026 a poor month to lock routing rules into config.

Where Claude Opus 5.5 Still Fails or Underperforms

Claude Opus 5.5 trails GPT-6 Astra on AutomationBench and Terminal-Bench-Science, routes most cybersecurity tasks to a weaker fallback model, can’t run without thinking, dropped forced tool use, outputs text only, and carries benchmark and safety claims that no third party had replicated as of September 22, 2026.

The concrete gaps, with Anthropic’s own numbers:

  • AutomationBench: 40.0% against Astra’s 41.4%, on a Zapier-run test where every safeguard intervention counted as a failure. Close enough to be noise, but a loss on the sheet.
  • Terminal-Bench-Science 0.1: 58.7% against Astra’s 64.6%, with a standard error of ±3.5 to 5 points per model. The widest deficit in the table, and the one most likely to hold up.
  • Cybersecurity work: most tasks reroute to Opus 4.8 until the Cyber Verification Program expands, so a security team buying Opus 5.5 today is often getting a two-generation-old model on the tasks it cares about.
  • Biology tasks: the classifier falls back to Opus 5 without enrolment in the Life Sciences Verification Program.
  • No thinking-off mode: every request pays for at least low-effort reasoning, which raises the cost and latency floor on trivial calls.
  • No forced tool use: tool_choice accepts only auto and none, so pipelines that relied on forcing a specific tool need a redesign.
  • Text out only: images go in, text comes out. No image, audio or file generation.

The alignment caveats come from Anthropic itself, which is to its credit. The launch post says the model often suspects it is being evaluated, and that building evaluations that reliably catch every failure before deployment remains unsolved. Read those two sentences together. A model that behaves well when it thinks it’s being watched tells you less about unattended production behaviour than the 85% containment figure implies.

Then there’s the replication gap. METR and Frontier Design tested the model pre-release, but neither had published findings as of launch day, and the UK AI Security Institute, which ran external cyber-range testing on Opus 5, has published nothing on Opus 5.5. Anthropic’s Transparency Hub hadn’t listed the new system card either. Every benchmark score and every safety number in circulation traces to one source.

One more thing Anthropic volunteers: the Fable 5.1 margins in its own table overstate the real-world gap. Anyone choosing Opus 5.5 over Fable on the strength of a 10-point Terminal-Bench lead should expect something smaller in practice.

How to Migrate Existing Claude Integrations to Opus 5.5

Migrating an existing Claude integration to Opus 5.5 means swapping the model ID to claude-opus-5-5, removing any request that disables thinking, replacing thinking budgets with the effort parameter, dropping forced tool choice, updating the computer-use tool, and re-baselining cost at the new medium default. Anthropic’s migration notes list four breaking changes from Opus 5, and each one returns a hard error instead of a silent downgrade, which makes the failures easy to find in staging and painful to find in production.

Work through the list in this order:

  1. Swap the model ID. Replace claude-opus-5 (or older) with claude-opus-5-5 in every config, environment variable and hard-coded string. Grep for the old ID across infrastructure repos too, since Bedrock and Vertex model names often live outside application code.
  2. Delete thinking: {"type": "disabled"}. Any request still carrying it returns a 400. There is no replacement flag; thinking is always on.
  3. Replace manual thinking budgets with effort. Five levels are available (low, medium, high, xhigh, max). If your Opus 5 integration tuned a token budget per call, map it to the nearest level and expect medium to cover most of what high did before, per the Opus 5.5 migration documentation.
  4. Remove forced tool use. tool_choice set to any or a named tool now errors. Rewrite the prompt to request the tool and keep auto, or handle the “no tool called” branch explicitly.
  5. Check for mid-thread model switches. Thinking blocks are bound to the model and conversation that produced them. A router that starts a thread on Sonnet and escalates to Opus 5.5 with the same message history needs to strip or regenerate those blocks.
  6. Move to computer_toolset_20260801. The legacy computer_20251124 tool is rejected on the Claude API and Google Cloud.
  7. Test preserved thinking if your account is new. For API accounts created on or after August 31, 2026, edits to Claude’s prior context are blocked. Any flow that rewrites earlier assistant turns before resending will break.
  8. Add handling for two new refusal paths. A biology classifier now runs alongside the cyber one, and a reasoning_extraction category refuses prompts that try to pull out internal reasoning verbatim. Log both so you can tell a safeguard fallback from a model failure.
  9. Re-baseline cost after the change. Run the same eval suite at medium effort and compare tokens per task, cache-hit ratio and wall-clock time against your Opus 5 baseline. Anthropic’s 40% figure assumes default settings and typical cache use; your number will differ.

The step most teams skip is number five. It fails only when a real user triggers an escalation path, which is exactly when you’d rather it didn’t.

Frequently Asked Questions About the Claude Opus 5.5 Launch

These are the questions people search most often about the Claude Opus 5.5 launch, each answered in a few sentences with the facts as of September 22, 2026.

When was Claude Opus 5.5 released?

Anthropic released Claude Opus 5.5 on September 22, 2026, as the first model in the Claude 5.5 family. It arrived two months after Claude Opus 5 (July 24, 2026) and about three weeks after Claude Fable 5.1 and Mythos 5.1, which launched around September 1 to 2, 2026. The model went live on the API, Claude Code, the Claude apps and all three major clouds the same day.

Is Claude Opus 5.5 better than Claude Fable 5.1?

On Anthropic’s own September 2026 benchmarks, yes on most rows: Opus 5.5 leads Fable 5.1 on Terminal-Bench 4.0 (66.4% against 55.8%), GDPval-AA v2.1 (1846 against 1735 Elo) and every other listed test. Anthropic adds that the real-world gap is narrower than those scores suggest, and Fable 5.1 remains its higher tier. The fair reading is near-parity at 40% of Fable’s price.

Is Opus 5.5 cheaper than Opus 5?

Yes. List prices are 20% lower ($4/$20 per million input/output tokens against $5/$25), cache reads are 60% lower ($0.20 against $0.50 per million), and Anthropic measures about 40% lower cost on typical workloads at default settings because the model also uses fewer tokens per task. Fast mode at $8/$40 is the only tier priced above standard rates.

What is the context window of Claude Opus 5.5?

1,000,000 tokens, roughly 555,000 words, with a maximum output of 128,000 tokens on a synchronous call and up to 300,000 tokens on the Batch API with a beta header. That matches Opus 5 and Fable 5.1 and sits just under GPT-6 Astra’s roughly 1.05M-token window. Input accepts text and images. Output is text only.

Can you turn off thinking in Opus 5.5?

No. Adaptive thinking is always on, and a request that sets thinking to disabled returns a 400 error. Control comes from the effort parameter instead, with five levels (low, medium, high, xhigh, max) and medium as the default. For short, simple calls where thinking overhead matters, Sonnet 5 or Haiku is the practical workaround.

Is Claude Opus 5.5 available on Claude Pro?

Yes. Claude Pro ($17 to $20 a month) and Claude Max (from $100 a month with 5x or 20x Pro usage) both include the latest Opus-tier model as of September 22, 2026, per Claude’s plans and pricing page. The launch also raised five-hour usage limits on Pro, Max, Team and seat-based Enterprise plans and added a rate-limit reset subscribers can save for later.

When are Claude Sonnet 5.5 and Haiku 5.5 coming?

Anthropic says both will follow “in the coming weeks” after the September 22, 2026 Opus 5.5 release, carrying many of the same performance, efficiency and safety changes. No exact dates are published. Until they ship, Sonnet 5 at $2/$10 per million tokens remains the current mid-tier option.

What is the knowledge cutoff for Claude Opus 5.5?

June 2026, for both reliable knowledge and training data, per Anthropic’s model documentation. That’s five months more recent than Claude Sonnet 5’s January 2026 cutoff. Anthropic hasn’t published cutoffs for Opus 5 or Fable 5.1 in the same comparison table.

AlphaCorp AIonline
Let's talk

Curious what AI could do for your business?

No jargon and no hard sell. Just a friendly look at where AI fits, and where it doesn't.

View Services

Is Claude Opus 5.5 safe, and who tested it?

Anthropic reports the best score of any Claude model on its roughly 2,000-scenario behavioural audit and about 85% fewer containment-boundary attempts than Opus 5. Pre-release testing came from Anthropic’s own red teams plus METR and Frontier Design. As of September 22, 2026, no third party had published an independent evaluation, so every safety claim rests on Anthropic’s system card.

Where to Start with Claude Opus 5.5

Start with your own eval suite, run against claude-opus-5-5 at the medium default, with prompt caching switched on before you read a single cost number. Skip the caching step and the 60% cache-read cut never shows up, which makes any comparison to Opus 5 worthless.

Then, in order:

  1. Pilot one long-horizon job (a migration, an audit, a research report) where the step count is high enough for the efficiency gains to register.
  2. Log tokens per task, turns per task and wall-clock time against your Opus 5 baseline. Anthropic’s 40% is an average at default settings. Yours will differ.
  3. If biology or security work is in scope, apply to the Life Sciences Verification Program now and watch for the Cyber Verification Program expansion.
  4. Hold routing rules loosely until Sonnet 5.5 and Haiku 5.5 land and until METR or the UK AI Security Institute publish something on this model.

The Opus 5.5 launch rewards teams that measure before they switch. If you’d rather have engineers who’ve already run this migration size it against your workload, talk to AlphaCorp AI about a scoped Opus 5.5 pilot.

Share
Newsletter · Weekly

Stay Ahead of AI

One email per week with the AI engineering insights, agent builds, and tools that actually matter.

No spamUnsubscribe anytimeFree forever

In every issue
  1. 01One agent build, taken apart step by step
  2. 02The tools that earned a place in our stack this week
  3. 03What broke in production, and what we changed

Written by Ignas Vaitukaitis, founder of AlphaCorp AI.

Wireframe cubes of circuitry linked by glowing strands above a dark circuit-board floor

Ready to Ship
Your AI System?

Book a free call and let's talk about what AI can do for your business. No sales pitch, just a real conversation.