← BlogBlog

Migrating to Claude Sonnet 5.5: 7 essential tips for a clean switch

Migrating to Claude Sonnet 5.5 looks like a one-line change, and for plenty of apps it nearly is: the price per token and the tokenizer are the same as Sonnet 5. But five settings that worked yesterday now return a 400, and a few prompt lines that used to help now make things worse. The seven tips below follow the order we would do the work in.

$2/$10Per MTok in / out, same as Sonnet 5
5Breaking changes from Sonnet 5
1MContext window, 128K max output
Sep 282026 release date

1. Claude Sonnet 5.5 keeps the price and changes the defaults

Features overview. Anthropic released Claude Sonnet 5.5 on September 28, 2026. The model overview lists the ID as claude-sonnet-5-5 with no date suffix (anthropic.claude-sonnet-5-5 on Bedrock), a 1M-token context window and 128K max output, or 300K on Message Batches with the output-300k-2026-03-24 beta header. Adaptive thinking is on by default, and effort runs from low to max, defaulting to high.

Prices are $2 / $10 per million input / output tokens and $0.20 for cache reads, identical to Sonnet 5. The pricing page also confirms that Sonnet 5's planned increase to $3 / $15 on September 1 will not happen, so staying on Sonnet 5 a little longer costs nothing extra.

Key differences. Per what's new in Sonnet 5.5, the same text produces the same token counts as on Sonnet 5. Coming from 4.x, the migration guide counts about 30% more tokens for the same text and about 2.5 times as many for a 2000×1500 image.

What changes, depending on where you start
PropertyFrom Sonnet 5From Sonnet 4.6 / 4.5 / Haiku 4.5
Tokens for the same textIdenticalAbout 30% more
Minimum cacheable prompt512 (was 1,024)512 (was 1,024; 4,096 on Haiku 4.5)
Thinking offdisabled is a 400; use between_toolsThinking now runs when the field is omitted
budget_tokens, sampling paramsNothing newReplace with effort; drop non-default temperature / top_p / top_k
On agentic benchmarks Sonnet 5.5 lands close to Opus 5.5Source: Anthropic's Claude Sonnet 5.5 announcement (retrieved Sep 29, 2026)
CursorBench 4.0 - Sonnet 5.5
55.5%
CursorBench 4.0 - Sonnet 5
34.1%
CursorBench 4.0 - Opus 5.5
57.8%
OSWorld 2.1 - Sonnet 5.5
80.1%
OSWorld 2.1 - Sonnet 5
57.0%
OSWorld 2.1 - Opus 5.5
81.8%

Anthropic's announcement says Sonnet 5.5 is 30%+ faster than Sonnet 5 and, at Low or Medium effort, beats Sonnet 5's best score "for about a tenth of the cost per task". Artificial Analysis looked at the other end of the dial. At max it scores 56 on their Intelligence Index, two points behind Opus 5.5, but used about 193K output tokens per task and cost $7.60 per task, roughly 50% more than Sonnet 5. Rajesh Beri's comparison of the same data has Opus 5.5 at xhigh reaching that 56 for $3.46.

So the per-task bill depends mostly on the effort level you pick, and cheaper is something you verify on your own traffic.

2. Prepare for migrating to Claude Sonnet 5.5 before you touch the model ID

Inventory what breaks. A plain text search across the repo finds nearly everything, as long as it covers config and wrapper-library settings, not only your own API calls.

What to search for
Search forOn Sonnet 5.5
"disabled" in a thinking object400
budget_tokens400, even with between_tools
tool_choice of any or tool400, on count_tokens too
Assistant prefill, non-default temperature / top_p / top_k400
computer_20251124400 on the Claude API and Google Cloud
Advisor model names400 for Opus 4.8, 4.7, 4.6, Sonnet 5, Sonnet 4.6
Code that edits messages, system or tools mid-conversation400 on newer accounts when old thinking blocks are replayed

Build the eval before the swap. Capture a baseline on the current model first, or every later comparison runs on memory. Anthropic's guide to building evals says to "mirror your real-world task distribution" and adds that more questions with automated grading beat fewer hand-graded ones. A hundred real requests from last week's logs beat twenty showcase prompts.

Let tools do part of it. In Claude Code, /claude-api migrate this project to claude-sonnet-5-5 runs the bundled migration skill, which swaps IDs, rewrites breaking parameters and prefills, and "produces a checklist of items that require manual verification". The guide's ant CLI samples won't help on Bedrock, which the Bedrock page says the CLI does not support.

3. Prompt engineering techniques for Sonnet 5.5 start with deleting lines

The Sonnet 5.5 prompting page opens with: "Existing Claude Sonnet 5 prompts should perform well without changes." The trouble is in the corrective lines teams piled up over earlier model generations.

What to remove. The best-practices page says to dial back anti-laziness prompting: where you wrote "CRITICAL: You MUST use this tool when...", plain "Use this tool when..." is enough. Drop "minimize tool calls" and "only use tools when strictly necessary", and drop "hold all findings for the final response", which suppresses the progress notes the model now writes between tool calls.

Two removals matter more than the rest. Asking the model to think less "doesn't reliably reduce its thinking", and with between_tools it makes the model "more likely" to leak internal XML tags into visible output, so lower the effort instead. And prompts asking the model to write out its reasoning now invite reasoning_extraction refusals.

What to add. The prompting page documents a few lines with measured effects.

Lines worth adding, with the effect Anthropic reports
AddWhereReported effect
"Think the problem through before you answer."JSON answersAt high, accuracy close to xhigh
The verification paragraph ("When you change code that can be run... run a real check...")Coding agents at lowSkipped checks become rare, no measurable quality change
The stop paragraph ("When the work ... is done and its checks pass, stop and report...")Long agent sessionsAbout a third off session cost at max, same quality
The search-first paragraph ("Use the search tool to check specifics...")Support and researchPrefers current sources over training knowledge

Testing and refining. Effort is now the main dial, and the level you ran on Sonnet 5 deserves a re-test rather than a copy. The effort docs say: "Start at high unless your workload is agentic or latency-sensitive. For agentic coding and multistep tool use, start at medium for well-specified tasks and move to high for harder or longer ones. For chat and other latency-sensitive work, start at medium or low." Then run your eval at two or three levels. We put every agent build through this sweep before a model change, and it is the step we would keep if we could keep only one.

4. Five settings from Sonnet 5 now return a 400

The release notes put it plainly: "Code written for Claude Sonnet 5 can break on Claude Sonnet 5.5 in five ways." Each fails loudly, in the first test run.

Setting, error, fix
Sonnet 5 code sendsSonnet 5.5 returnsFix
thinking: {"type": "disabled"}400, pointing to between_toolsAdaptive thinking at low, or between_tools at high or below
tool_choice of any / tool400: tool_choice: type "tool" and "any" are not supported for this model.auto plus strict: true, or structured outputs
Thinking blocks replayed after a history edit400 on accounts created from Aug 31, 2026Append-only history; drop_block (adaptive only)
computer_20251124400 on Claude API and Google Cloudcomputer_toolset_20260801; Bedrock keeps the old tool
Advisor on Opus 4.8 / 4.7 / 4.6, Sonnet 5 or 4.6400Opus 5 / 5.5, Sonnet 5.5, Fable or Mythos 5 / 5.1

Thinking off has rules. between_tools works only at high effort or below, allows no other field inside thinking, and rejects per-message effort changes. Before using it, try adaptive thinking at low, where the model "skips thinking on most simple requests". Addy Osmani's launch post notes that with between_tools "total response time is the same or faster". Measure both.

Check your wrappers. In a kinby issue, "Recap fails on claude-sonnet-5-5 because it forces tool_choice" - the code never set tool_choice itself. LangChain's with_structured_output() did, and the fix was method="json_schema". Since auto does not guarantee a call, check that it happened. On Bedrock, structured outputs and strict tool use aren't available for Sonnet 5.5, so validate in your own code.

Thinking blocks are bound. Sonnet 5.5 reads blocks from Sonnet 5, so live conversations carry over, but no other model reads Sonnet 5.5's blocks; they are silently dropped. The preserved thinking docs add a history check for accounts created on or after August 31, 2026 on the Claude API, Bedrock and Google Cloud: replaying a block after editing the system prompt, tools or earlier messages is a 400. Beri reports the Turnstone framework's automatic compaction failing on exactly this.

Legacy considerations. From Sonnet 4.6 or earlier, budget_tokens must become an effort level. From Sonnet 4.5 and Haiku 4.5, assistant prefills return a 400. Haiku 4.5 users may want to wait: VentureBeat reports Haiku 5.5 is due within weeks.

5. The new Claude Sonnet features pay off in cache hits and latency

Per-message effort (beta mid-conversation-output-config-2026-07-01) lets one hard question run at high inside a low session. That matters because changing top-level effort between requests "invalidates prompt caching", per the effort docs. It needs adaptive thinking.

Mid-conversation system messages and inline tools let you add an instruction or a tool by appending instead of editing the system prompt, which keeps earlier thinking blocks valid. One caution from the prompting page: "Never put user text inside a tool_result block".

Compaction on demand (beta compact-2026-09-04, since September 14) gives long agent sessions a controlled point to shrink context instead of hand-editing history.

Progress notes changed shape. Text between tool calls now arrives in thinking blocks, empty by default, so a UI that renders only text goes quiet with no error. Set display: "updates" or "summarized" with adaptive thinking.

Smaller things add up too. The tool-use system prompt is 286 tokens against 354 on Sonnet 5, and short prompts now cache from 512 tokens. Lovable told VentureBeat it saw one-third fewer tool calls, though that is their workload and not yours.

6. After the switch, measure cost per completed task

With the same per-token price, the number that moves is cost per completed task: every token spent, thinking included, divided by the tasks that passed your eval.

Post-migration metrics
MetricWhy it movesWatch for
Cost per completed taskEffort drives output tokensCompare with the baseline per effort level
Tool calls per taskOften fewer on Sonnet 5.5Zero calls on a tool you need
Time to first token (p50, p95)Thinking runs before the replylow skips it on most simple requests
Refusals by categorycyber, bio, frontier_llm, reasoning_extraction, general_harmsBranch on stop_reason before reading content

Refusals cost money again. A refusal is an HTTP 200 with stop_reason: "refusal". Since September 24, pre-output refusals in bio, frontier_llm and reasoning_extraction are billed. Server-side fallback retries only cyber and frontier_llm on Sonnet 5; the rest is yours to handle.

Silent regressions. Some failures never throw. A router switch or fallback runs the next turns without Sonnet 5.5's reasoning. A stop_reason: "max_tokens" with JSON that happens to parse is still a failure; for agentic coding the docs recommend max_tokens of 128,000 with streaming.

Close the loop. Assert that response.model starts with claude-sonnet-5-5 in your smoke test. Mind the prefix: it also starts with claude-sonnet-5, so an old Sonnet 5 check passes on both. Unmeasured changes are the thread running through why AI deployments fail, and a model swap is one.

7. Put the next model swap on a schedule

The useful output of this migration is a procedure you can repeat. The deprecation page promises "at least 60 days' notice" before retirement and lists the earliest dates.

Earliest retirement dates, as listed on September 29, 2026
ModelRetirement
claude-sonnet-5-5Not sooner than Sep 28, 2027
claude-sonnet-5Not sooner than Jun 30, 2027
claude-sonnet-4-6Not sooner than Feb 17, 2027
claude-sonnet-4-5-20250929Not sooner than Sep 29, 2026 (date reached; no retirement notice as of Sep 29, 2026)
claude-haiku-4-5-20251001Not sooner than Oct 15, 2026

Bookmark the release notes, the deprecation page and the per-model prompting page. The best-practices page adds the rule behind every tip here: re-check a prompt "against your own evals before applying it to another" model. The changelog below records each time we re-check the dates here.

It is also a good moment to ask whether a route needs a model at all; agents vs workflows covers that trade-off.

The Sonnet 5.5 migration checklist you can run alone

Capture a baseline on the current model: pass rate, tokens, tool calls, latency.

Search the repo using the table in tip 2, wrappers and config included.

Swap the ID to claude-sonnet-5-5 (anthropic.claude-sonnet-5-5 on Bedrock).

Replace disabled thinking with adaptive at low, or between_tools at high or below.

Replace forced `tool_choice` with auto plus strict: true or structured outputs, and check the call happened.

Clean 4.x leftovers: budget_tokens, sampling parameters, prefills.

Keep history append-only and pass thinking blocks back unchanged.

Update computer use and advisor pairings.

Handle refusals by category before reading content.

Edit prompts: remove think-less and write-your-reasoning lines, add the documented paragraphs.

Sweep effort at two or three levels and set it explicitly.

Re-baseline cost per completed task and p95 latency, and assert the response.model prefix.

If you want a second pair of eyes on an agent running on Claude, that's part of what we look at when you book a free process audit. The checklist works fine without us, and the first two items are worth doing this week even if the switch waits a month.

FAQ

Do I have to change my prompts when migrating to Claude Sonnet 5.5?

Usually not the core. Remove lines asking the model to think less or write out its reasoning, and add the documented paragraphs only where your eval shows they help.

Is Sonnet 5.5 cheaper than Sonnet 5?

Per token, no: both cost $2 / $10 per million. Per task it depends on effort. Anthropic reports that at low or medium effort it beats Sonnet 5's best score for about a tenth of the cost per task; Artificial Analysis measured about 50% more per task at max.

Can I keep thinking off?

Not with disabled, which returns a 400. The closest is between_tools, limited to high effort or below with no other thinking fields. Try adaptive thinking at low first.

What changes on Amazon Bedrock?

The ID is anthropic.claude-sonnet-5-5. Structured outputs, strict tool use, the advisor tool, Message Batches, server-side fallback and the ant CLI are not available there, while computer_20251124 is still accepted.

Changelog
  • 29 September 2026Published.
Free process audit

See what this would look like in your operations.

Get in touch

30 minutes · we map your 3 best automation opportunities · no obligation