Claude Sonnet 5.5 vs Sonnet 5: 7 Real Differences
Claude Sonnet 5.5 vs Sonnet 5: pricing, Terminal-Bench 4.0 scores, agent reliability, migration gotchas, and whether to upgrade.
Claude Sonnet 5.5 vs Sonnet 5: pricing, Terminal-Bench 4.0 scores, agent reliability, migration gotchas, and whether to upgrade.

Anthropic quietly dropped Sonnet 5.5 on September 28, 2026 as the second model in the Claude 5.5 family, following Opus 5.5 by about a week. The jump from Sonnet 5 is bigger than the point-release label suggests — Anthropic frames it as "a clear upgrade" with 30%+ faster inference and up to 30% lower cost for most workloads.
The short answer: yes, upgrade. The longer answer involves one benchmark number that will make you double-check it, a pricing structure that quietly stayed the same per token, and a few migration gotchas around tool use. Let's get into it.
Claude Sonnet 5.5 is a drop-in upgrade over Claude Sonnet 5 with large gains in agentic coding, meaningful improvements in long-context reasoning and vision, and faster inference — all at the same list price per token. The only real friction is a migration quirk: if you were running Sonnet with thinking off, you'll need to flip to the new between_tools setting before you move.
If you're building an agent, upgrade this week. If you're on a hot code path with tight prompt-response SLAs, you still want to A/B test, but you're mostly A/B testing to confirm the win, not to find a regression. That's the gist.
| Attribute | Claude Sonnet 5 | Claude Sonnet 5.5 |
|---|---|---|
| Release | Earlier 2026 | September 28, 2026 |
| Context window | 200K tokens | 200K tokens (1M in beta) |
| Input price | $2 / MTok | $2 / MTok |
| Output price | $10 / MTok | $10 / MTok |
| Cache reads | Standard discount | $0.20 / MTok |
| Extended thinking | Yes | Yes, with interleaved tool calls |
| Vision | Yes | Yes, better charts and multi-column PDFs |
| Tool use reliability | Strong | Noticeably stronger |
| Terminal-Bench 4.0 | 10.3% | 70.6% |
| Best use case | Everyday coding, writing | Long-horizon agents, SWE tasks |
Two things stand out. List pricing didn't move, which is the quiet story. And that Terminal-Bench 4.0 jump — from 10.3% to 70.6% — isn't a typo. It's Anthropic's own number on an agentic coding evaluation, and it's one of the largest version-over-version deltas we've seen from a point release.
Anthropic published a dedicated announcement page for Sonnet 5.5, and the Anthropic news feed and Claude Code docs give enough signal to piece together the real changes. Three categories matter.
Sonnet 5.5 earns its half-point and then some. On Terminal-Bench 4.0, an agentic coding evaluation, Anthropic reports Sonnet 5.5 at 70.6% versus Sonnet 5 at 10.3%. That's not a tuning win — that's a new capability tier.
On SWE-bench Verified specifically, the Sonnet line has consistently punched above its weight. The SWE-bench leaderboard shows the top of the chart dominated by frontier reasoning models from Anthropic and OpenAI, and Sonnet 5.5 continues that climb.
What matters more than any single benchmark is the failure-mode shift. The old Sonnet 5 would sometimes spiral when a shell command returned an unexpected error, narrating around it instead of actually recovering. 5.5 recovers. Real agent work is 40 steps of fix-the-last-step, and the model's willingness to stop, look at the error, and try again is what makes those loops converge.
If you've shipped anything with function calling, you know the pain: the model decides to narrate a tool call instead of actually emitting it, or passes a string where you asked for an int. Sonnet 5 was already pretty solid here. 5.5 is better again, especially around parallel tool calls and interleaved thinking (where the model can think between tool call results, mid-conversation).
One migration note worth flagging: if you ran Sonnet with thinking off, Sonnet 5.5 introduces a new between_tools setting that controls up-front thinking separately from inter-tool thinking. You'll need to opt into that before switching.
Sonnet 5 could read a screenshot. Sonnet 5.5 reads a 50-page PDF with charts, tables, and footnotes, and gives you back something structured. It's also the first Sonnet model to beat Pokémon Red working only from screenshots, which is a stronger agentic vision benchmark than it sounds — long-horizon play from pixels is hard.
If you're doing document extraction, this is the single biggest reason to upgrade.
Both models sit at the same list price: $2 per million input tokens and $10 per million output tokens. The headline number didn't move. For reference:
| Model | Input ($/MTok) | Output ($/MTok) |
|---|---|---|
| Claude Sonnet 5 | $2 | $10 |
| Claude Sonnet 5.5 | $2 | $10 |
| Claude Opus 5.5 | $4 | $20 |
| GPT-4o | $2.50 | $10 |
So Sonnet 5.5 is price-matched to its predecessor and roughly in line with GPT-4o on input, cheaper on output. But the real cost story is efficiency: Anthropic says Sonnet 5.5 completes typical work with roughly 30% fewer tokens, which means your actual bill drops even though the per-token price is flat.
The per-token price didn't move, but Sonnet 5.5 is still cheaper to run than Sonnet 5. Fewer tokens, same rate.
Opus 5.5 is 2x the output price of Sonnet 5.5. If you were routing everything to Opus for safety, Sonnet 5.5 is the pragmatic sweet spot for most jobs now. Prompt caching helps even more — cache reads sit at $0.20 per million tokens, and most agentic workloads lean heavily on cache hits. If you're not using prompt caching with Sonnet 5.5, you're leaving a lot of money on the table.
Anthropic doesn't publish every standard benchmark consistently, but the ones in the official announcement give a clear picture. The LMSYS Chatbot Arena tracks Elo over time, and the Anthropic family has been climbing steadily. Sonnet 5.5 isn't tracked separately in all slices yet, but community estimates place it above Sonnet 5 by a meaningful margin.
For agentic coding, Terminal-Bench 4.0 is the headline: 70.6% for Sonnet 5.5 versus 10.3% for Sonnet 5. On general code generation, HumanEval is approaching saturation across the board, so SWE-bench Verified and Terminal-Bench are the better directional signals. Sonnet 5.5 is meaningfully stronger on tasks that require editing multiple files or understanding repo structure.
On GDPval-AA — a real-world work evaluation across occupations that Artificial Analysis runs — Anthropic reports Sonnet 5.5 landing two points below Opus 5.5. That's close enough that for a lot of knowledge work, Sonnet 5.5 is now a reasonable Opus-tier substitute at half the output price.
For PhD-level math or GPQA-Diamond-style questions, Opus is still the pick. Sonnet 5.5 isn't a dedicated reasoning model. If your workload is dominated by hard science or math, you'll feel the ceiling.
Sonnet 5 could technically handle 200K tokens, but quality degraded past roughly 100K. Sonnet 5.5 holds up noticeably better at the 150K-180K range based on community needle-in-haystack testing. If you're doing codebase-level analysis or long document QA, this alone justifies the upgrade.
Both models support extended thinking (Anthropic's name for inference-time reasoning). The difference: Sonnet 5.5 can interleave thinking blocks with tool calls, meaning the model can think, call a tool, think again, call another tool. Sonnet 5 could only think once before responding.
This is genuinely useful for agents. It means the model can reflect on tool output mid-conversation rather than committing to a plan upfront.
The computer use beta works on both, but Sonnet 5.5 is substantially better at it. Success rates on OSWorld-style tasks are up by a decent margin per Anthropic's own evals. Still not production-ready for most workflows, but closer.
Same API, cache reads at $0.20 per million tokens. If you have a 50K-token system prompt (common for coding agents), cache it. The read discount makes repeat queries cheap.
Native PDF support exists on both, but Sonnet 5.5 handles multi-column layouts and dense charts significantly better. If you've been post-processing PDFs with OCR before sending them to Claude, you can probably stop.
Anthropic reports Sonnet 5.5 running 30%+ faster than Sonnet 5 overall. That's a real win for both streaming chat UIs and agent loops — faster time-to-first-token on short prompts, and faster end-to-end on longer tasks.
Pick Claude Sonnet 5.5 if:
Stick with Claude Sonnet 5 if:
Honestly, the second list is short. For 95% of use cases, 5.5 is a straight upgrade. For a related head-to-head, see Claude Sonnet 4.6 vs GPT-4o.
Sonnet 5.5 is the pick. The tool use reliability, recovery from failed commands, and the Terminal-Bench 4.0 numbers all point the same direction. Pair it with Claude Code or a custom agent and you'll see fewer frustrating infinite loops. If you're also weighing editor choices, Claude Code vs Cursor vs Copilot covers that angle.
Either works. If your bot is doing simple intent classification and retrieval, Sonnet 5 is fine and the migration cost isn't worth it yet. If it needs to handle multi-step issues (refund + account update + email draft), 5.5's tool orchestration is worth the swap.
Sonnet 5.5, no question. The vision improvements for PDFs and charts are the single clearest upgrade in this release.
Sonnet 5.5. The quality at 150K+ tokens is noticeably better.
Both models are good at this and the differences are minor. If you've tuned a specific voice against Sonnet 5, don't rock the boat.
Before flipping the model string in production:
claude-sonnet-5-5between_tools settingSonnet 5.5 is the kind of release Anthropic does well: no hype, no new buzzwords, just the model getting better at the stuff you were already using it for — with one benchmark result (Terminal-Bench 4.0) that quietly reframes what Sonnet-tier agentic coding looks like. The flat per-token pricing combined with 30% fewer tokens for most work makes the upgrade decision close to zero-risk for most teams.
But don't expect a transformation of what Sonnet is. If Sonnet 5 wasn't working for you because you needed harder reasoning, 5.5 still won't fix that — Opus 4.7 vs GPT-5.5 covers the higher tier. For everyone else: flip the model string, your agent runs more reliably, your bill drops a bit, and you move on with your life. That's a good day.
Mostly yes. The API surface is identical — update the model string to `claude-sonnet-5-5` and your tool schemas, system prompts, and streaming code work unchanged. One migration gotcha: if you ran Sonnet with thinking off, you'll need to opt into the new `between_tools` setting before switching. Also re-test any prompts that worked around specific Sonnet 5 quirks.
Yes, but it's gated behind a beta flag and tier-based access, same as the Sonnet 5 1M beta. You'll need to be on a qualifying usage tier (or request access via your account manager) and pass the appropriate beta header in your API request. Pricing for prompts above 200K tokens uses a separate rate card.
For agentic coding tasks like repo-level edits, multi-file refactors, and tool-driven workflows, Sonnet 5.5 is the stronger choice based on Terminal-Bench 4.0 and SWE-bench Verified results. For one-shot code snippets or quick Q&A, GPT-4o is competitive and marginally cheaper on input. The gap widens as task complexity and tool-use requirements increase.
Mostly, but expect to re-tune anything that depended on Sonnet 5 tool-call quirks or thinking patterns. Sonnet 5.5 is less likely to narrate instead of calling tools, which can change the behavior of prompts that worked around that bug. Run a regression test suite before flipping production traffic, and watch tokens-per-request to confirm the efficiency gain lands.
Use Opus 5.5 when the task demands harder reasoning: PhD-level science questions, complex math, or ambiguous multi-step planning where mistakes compound. Opus 5.5 costs 2x the output rate of Sonnet 5.5 ($20 vs $10 per MTok), so reserve it for the hard chunk of your workload and route everything else to Sonnet 5.5. On GDPval-AA, Sonnet 5.5 is only two points behind Opus 5.5.