1. Claude Sonnet 5.5 keeps the price and changes the defaults
Features overview. Anthropic released Claude Sonnet 5.5 on September 28, 2026. The model overview lists the ID as claude-sonnet-5-5 with no date suffix (anthropic.claude-sonnet-5-5 on Bedrock), a 1M-token context window and 128K max output, or 300K on Message Batches with the output-300k-2026-03-24 beta header. Adaptive thinking is on by default, and effort runs from low to max, defaulting to high.
Prices are $2 / $10 per million input / output tokens and $0.20 for cache reads, identical to Sonnet 5. The pricing page also confirms that Sonnet 5's planned increase to $3 / $15 on September 1 will not happen, so staying on Sonnet 5 a little longer costs nothing extra.
Key differences. Per what's new in Sonnet 5.5, the same text produces the same token counts as on Sonnet 5. Coming from 4.x, the migration guide counts about 30% more tokens for the same text and about 2.5 times as many for a 2000×1500 image.
| Property | From Sonnet 5 | From Sonnet 4.6 / 4.5 / Haiku 4.5 |
|---|---|---|
| Tokens for the same text | Identical | About 30% more |
| Minimum cacheable prompt | 512 (was 1,024) | 512 (was 1,024; 4,096 on Haiku 4.5) |
| Thinking off | disabled is a 400; use between_tools | Thinking now runs when the field is omitted |
budget_tokens, sampling params | Nothing new | Replace with effort; drop non-default temperature / top_p / top_k |
Anthropic's announcement says Sonnet 5.5 is 30%+ faster than Sonnet 5 and, at Low or Medium effort, beats Sonnet 5's best score "for about a tenth of the cost per task". Artificial Analysis looked at the other end of the dial. At max it scores 56 on their Intelligence Index, two points behind Opus 5.5, but used about 193K output tokens per task and cost $7.60 per task, roughly 50% more than Sonnet 5. Rajesh Beri's comparison of the same data has Opus 5.5 at xhigh reaching that 56 for $3.46.
So the per-task bill depends mostly on the effort level you pick, and cheaper is something you verify on your own traffic.
2. Prepare for migrating to Claude Sonnet 5.5 before you touch the model ID
Inventory what breaks. A plain text search across the repo finds nearly everything, as long as it covers config and wrapper-library settings, not only your own API calls.
| Search for | On Sonnet 5.5 |
|---|---|
"disabled" in a thinking object | 400 |
budget_tokens | 400, even with between_tools |
tool_choice of any or tool | 400, on count_tokens too |
Assistant prefill, non-default temperature / top_p / top_k | 400 |
computer_20251124 | 400 on the Claude API and Google Cloud |
| Advisor model names | 400 for Opus 4.8, 4.7, 4.6, Sonnet 5, Sonnet 4.6 |
Code that edits messages, system or tools mid-conversation | 400 on newer accounts when old thinking blocks are replayed |
Build the eval before the swap. Capture a baseline on the current model first, or every later comparison runs on memory. Anthropic's guide to building evals says to "mirror your real-world task distribution" and adds that more questions with automated grading beat fewer hand-graded ones. A hundred real requests from last week's logs beat twenty showcase prompts.
Let tools do part of it. In Claude Code, /claude-api migrate this project to claude-sonnet-5-5 runs the bundled migration skill, which swaps IDs, rewrites breaking parameters and prefills, and "produces a checklist of items that require manual verification". The guide's ant CLI samples won't help on Bedrock, which the Bedrock page says the CLI does not support.
3. Prompt engineering techniques for Sonnet 5.5 start with deleting lines
The Sonnet 5.5 prompting page opens with: "Existing Claude Sonnet 5 prompts should perform well without changes." The trouble is in the corrective lines teams piled up over earlier model generations.
What to remove. The best-practices page says to dial back anti-laziness prompting: where you wrote "CRITICAL: You MUST use this tool when...", plain "Use this tool when..." is enough. Drop "minimize tool calls" and "only use tools when strictly necessary", and drop "hold all findings for the final response", which suppresses the progress notes the model now writes between tool calls.
Two removals matter more than the rest. Asking the model to think less "doesn't reliably reduce its thinking", and with between_tools it makes the model "more likely" to leak internal XML tags into visible output, so lower the effort instead. And prompts asking the model to write out its reasoning now invite reasoning_extraction refusals.
What to add. The prompting page documents a few lines with measured effects.
| Add | Where | Reported effect |
|---|---|---|
| "Think the problem through before you answer." | JSON answers | At high, accuracy close to xhigh |
| The verification paragraph ("When you change code that can be run... run a real check...") | Coding agents at low | Skipped checks become rare, no measurable quality change |
| The stop paragraph ("When the work ... is done and its checks pass, stop and report...") | Long agent sessions | About a third off session cost at max, same quality |
| The search-first paragraph ("Use the search tool to check specifics...") | Support and research | Prefers current sources over training knowledge |
Testing and refining. Effort is now the main dial, and the level you ran on Sonnet 5 deserves a re-test rather than a copy. The effort docs say: "Start at high unless your workload is agentic or latency-sensitive. For agentic coding and multistep tool use, start at medium for well-specified tasks and move to high for harder or longer ones. For chat and other latency-sensitive work, start at medium or low." Then run your eval at two or three levels. We put every agent build through this sweep before a model change, and it is the step we would keep if we could keep only one.
4. Five settings from Sonnet 5 now return a 400
The release notes put it plainly: "Code written for Claude Sonnet 5 can break on Claude Sonnet 5.5 in five ways." Each fails loudly, in the first test run.
| Sonnet 5 code sends | Sonnet 5.5 returns | Fix |
|---|---|---|
thinking: {"type": "disabled"} | 400, pointing to between_tools | Adaptive thinking at low, or between_tools at high or below |
tool_choice of any / tool | 400: tool_choice: type "tool" and "any" are not supported for this model. | auto plus strict: true, or structured outputs |
| Thinking blocks replayed after a history edit | 400 on accounts created from Aug 31, 2026 | Append-only history; drop_block (adaptive only) |
computer_20251124 | 400 on Claude API and Google Cloud | computer_toolset_20260801; Bedrock keeps the old tool |
| Advisor on Opus 4.8 / 4.7 / 4.6, Sonnet 5 or 4.6 | 400 | Opus 5 / 5.5, Sonnet 5.5, Fable or Mythos 5 / 5.1 |
Thinking off has rules. between_tools works only at high effort or below, allows no other field inside thinking, and rejects per-message effort changes. Before using it, try adaptive thinking at low, where the model "skips thinking on most simple requests". Addy Osmani's launch post notes that with between_tools "total response time is the same or faster". Measure both.
Check your wrappers. In a kinby issue, "Recap fails on claude-sonnet-5-5 because it forces tool_choice" - the code never set tool_choice itself. LangChain's with_structured_output() did, and the fix was method="json_schema". Since auto does not guarantee a call, check that it happened. On Bedrock, structured outputs and strict tool use aren't available for Sonnet 5.5, so validate in your own code.
Thinking blocks are bound. Sonnet 5.5 reads blocks from Sonnet 5, so live conversations carry over, but no other model reads Sonnet 5.5's blocks; they are silently dropped. The preserved thinking docs add a history check for accounts created on or after August 31, 2026 on the Claude API, Bedrock and Google Cloud: replaying a block after editing the system prompt, tools or earlier messages is a 400. Beri reports the Turnstone framework's automatic compaction failing on exactly this.
Legacy considerations. From Sonnet 4.6 or earlier, budget_tokens must become an effort level. From Sonnet 4.5 and Haiku 4.5, assistant prefills return a 400. Haiku 4.5 users may want to wait: VentureBeat reports Haiku 5.5 is due within weeks.
5. The new Claude Sonnet features pay off in cache hits and latency
Per-message effort (beta mid-conversation-output-config-2026-07-01) lets one hard question run at high inside a low session. That matters because changing top-level effort between requests "invalidates prompt caching", per the effort docs. It needs adaptive thinking.
Mid-conversation system messages and inline tools let you add an instruction or a tool by appending instead of editing the system prompt, which keeps earlier thinking blocks valid. One caution from the prompting page: "Never put user text inside a tool_result block".
Compaction on demand (beta compact-2026-09-04, since September 14) gives long agent sessions a controlled point to shrink context instead of hand-editing history.
Progress notes changed shape. Text between tool calls now arrives in thinking blocks, empty by default, so a UI that renders only text goes quiet with no error. Set display: "updates" or "summarized" with adaptive thinking.
Smaller things add up too. The tool-use system prompt is 286 tokens against 354 on Sonnet 5, and short prompts now cache from 512 tokens. Lovable told VentureBeat it saw one-third fewer tool calls, though that is their workload and not yours.
6. After the switch, measure cost per completed task
With the same per-token price, the number that moves is cost per completed task: every token spent, thinking included, divided by the tasks that passed your eval.
| Metric | Why it moves | Watch for |
|---|---|---|
| Cost per completed task | Effort drives output tokens | Compare with the baseline per effort level |
| Tool calls per task | Often fewer on Sonnet 5.5 | Zero calls on a tool you need |
| Time to first token (p50, p95) | Thinking runs before the reply | low skips it on most simple requests |
| Refusals by category | cyber, bio, frontier_llm, reasoning_extraction, general_harms | Branch on stop_reason before reading content |
Refusals cost money again. A refusal is an HTTP 200 with stop_reason: "refusal". Since September 24, pre-output refusals in bio, frontier_llm and reasoning_extraction are billed. Server-side fallback retries only cyber and frontier_llm on Sonnet 5; the rest is yours to handle.
Silent regressions. Some failures never throw. A router switch or fallback runs the next turns without Sonnet 5.5's reasoning. A stop_reason: "max_tokens" with JSON that happens to parse is still a failure; for agentic coding the docs recommend max_tokens of 128,000 with streaming.
Close the loop. Assert that response.model starts with claude-sonnet-5-5 in your smoke test. Mind the prefix: it also starts with claude-sonnet-5, so an old Sonnet 5 check passes on both. Unmeasured changes are the thread running through why AI deployments fail, and a model swap is one.
7. Put the next model swap on a schedule
The useful output of this migration is a procedure you can repeat. The deprecation page promises "at least 60 days' notice" before retirement and lists the earliest dates.
| Model | Retirement |
|---|---|
claude-sonnet-5-5 | Not sooner than Sep 28, 2027 |
claude-sonnet-5 | Not sooner than Jun 30, 2027 |
claude-sonnet-4-6 | Not sooner than Feb 17, 2027 |
claude-sonnet-4-5-20250929 | Not sooner than Sep 29, 2026 (date reached; no retirement notice as of Sep 29, 2026) |
claude-haiku-4-5-20251001 | Not sooner than Oct 15, 2026 |
Bookmark the release notes, the deprecation page and the per-model prompting page. The best-practices page adds the rule behind every tip here: re-check a prompt "against your own evals before applying it to another" model. The changelog below records each time we re-check the dates here.
It is also a good moment to ask whether a route needs a model at all; agents vs workflows covers that trade-off.
The Sonnet 5.5 migration checklist you can run alone
Capture a baseline on the current model: pass rate, tokens, tool calls, latency.
Search the repo using the table in tip 2, wrappers and config included.
Swap the ID to claude-sonnet-5-5 (anthropic.claude-sonnet-5-5 on Bedrock).
Replace disabled thinking with adaptive at low, or between_tools at high or below.
Replace forced `tool_choice` with auto plus strict: true or structured outputs, and check the call happened.
Clean 4.x leftovers: budget_tokens, sampling parameters, prefills.
Keep history append-only and pass thinking blocks back unchanged.
Update computer use and advisor pairings.
Handle refusals by category before reading content.
Edit prompts: remove think-less and write-your-reasoning lines, add the documented paragraphs.
Sweep effort at two or three levels and set it explicitly.
Re-baseline cost per completed task and p95 latency, and assert the response.model prefix.
If you want a second pair of eyes on an agent running on Claude, that's part of what we look at when you book a free process audit. The checklist works fine without us, and the first two items are worth doing this week even if the switch waits a month.
FAQ
Do I have to change my prompts when migrating to Claude Sonnet 5.5?
Usually not the core. Remove lines asking the model to think less or write out its reasoning, and add the documented paragraphs only where your eval shows they help.
Is Sonnet 5.5 cheaper than Sonnet 5?
Per token, no: both cost $2 / $10 per million. Per task it depends on effort. Anthropic reports that at low or medium effort it beats Sonnet 5's best score for about a tenth of the cost per task; Artificial Analysis measured about 50% more per task at max.
Can I keep thinking off?
Not with disabled, which returns a 400. The closest is between_tools, limited to high effort or below with no other thinking fields. Try adaptive thinking at low first.
What changes on Amazon Bedrock?
The ID is anthropic.claude-sonnet-5-5. Structured outputs, strict tool use, the advisor tool, Message Batches, server-side fallback and the ant CLI are not available there, while computer_20251124 is still accepted.
- 29 September 2026Published.