Back to list
Sep 29, 2026
749
0
0
ClaudeNEW

Claude Sonnet 5.5: 30% Faster, Up to 30% Cheaper per Task

Claude Sonnet 5.5 keeps Sonnet 5 pricing, runs 30%+ faster, and cuts per-task cost up to 30%, per Anthropic; Opus 5.5 still leads most benchmarks.

#Claude#Anthropic#Claude Sonnet 5.5#LLM#Agentic Coding
Claude Sonnet 5.5: 30% Faster, Up to 30% Cheaper per Task
AI Summary

Claude Sonnet 5.5 keeps Sonnet 5 pricing, runs 30%+ faster, and cuts per-task cost up to 30%, per Anthropic; Opus 5.5 still leads most benchmarks.

Introduction

Anthropic released Claude Sonnet 5.5 on September 28, 2026, describing it as the second model in the Claude 5.5 family after Opus 5.5. According to Anthropic's announcement page, it is "a clear upgrade over Claude Sonnet 5," runs 30%+ faster, and costs up to 30% less for most work. Per-token prices are unchanged from Sonnet 5 at $2 per million input tokens and $10 per million output tokens, so the savings come from the model needing fewer tokens per task rather than from a list-price cut.

Anthropic positions Sonnet 5.5 as a faster, lower-cost complement to Opus 5.5. Where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is described as strongest at well-scoped everyday tasks, bug fixing, and polished documents, slides, and spreadsheets. TechCrunch, which covered the launch the same day, noted that Sonnet 5 was announced about three months earlier and that speed is the main selling point this time. For context on the larger sibling, see our earlier coverage of Claude Opus 5.5.

Feature Overview

1. Agentic coding gains. Anthropic reports 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, against 10.3% for Sonnet 5. On FrontierCode 1.1, which measures whether an agent's code changes would be merged, Sonnet 5.5 scores 52.1% at Xhigh effort. Anthropic says that at High effort it scores 10 points higher than Sonnet 5 at the same setting for about one fifteenth of the cost per task.

2. Speed and token efficiency. Anthropic says output generation is 30%+ faster than Sonnet 5, making this its fastest Sonnet model to date. In the company's testing it costs up to 30% less per task, and early testers observed that it batched tool calls together more than Sonnet 5, leading to fewer steps.

3. Knowledge work and computer use. On GDPval-AA v2.1, which tests real-world tasks across 44 occupations and nine industries, Sonnet 5.5 scores 1844 Elo versus 1846 for Opus 5.5 and 1449 for Sonnet 5. Anthropic also says it is the first Sonnet model to beat Pokémon Red working only from screenshots.

4. Effort control. The default effort is Medium in Claude Code and the Claude apps, and High on the Claude Platform. Anthropic says Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score on several benchmarks for about a tenth of the cost per task.

5. New safeguards. Because its cyber capabilities are comparable to Opus 5's, Anthropic says this is the first Sonnet model to launch with cyber safeguards and fallbacks. It is also the first Sonnet with classifiers that prevent reasoning extraction, and it expands "preserved thinking."

Published benchmarks (Anthropic's figures)

BenchmarkSonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Terminal-Bench 4.070.6%10.3%66.4% (Xhigh)not reported
FrontierCode 1.1 (Main)52.1% (Xhigh)42.4%54.4%49.3%
CursorBench 4.055.5%34.1%57.8%not reported
GDPval-AA v2.1 (Elo)1844144918461487
AA-Briefcase v1.1 (Elo)1811135918221483
Humanity's Last Exam (with tools)64.5%54.9%67.7%not reported
OSWorld 2.1 (partial)80.1%57.0%81.8%not reported
Chartography (no tools)61.6%15.6%64.4%53.6%

Read closely, the table does not show Sonnet 5.5 overtaking Opus 5.5 broadly. It leads Opus 5.5 only on Terminal-Bench 4.0, where the Opus figure is reported at Xhigh effort; Opus 5.5 is ahead on FrontierCode, CursorBench, GDPval-AA, AA-Briefcase, Humanity's Last Exam, and OSWorld. TechCrunch wrote that Anthropic's benchmarks show Sonnet 5.5 "performing better than Opus 5.5 on agentic coding," but that holds for one of the three coding rows. Anthropic itself says Opus 5.5 "remains clearly stronger at complex, open-ended work requiring sustained judgment." Two footnotes matter as well. Sonnet 5.5 scores lower at Max effort (46.2%) than at Xhigh on FrontierCode, which Anthropic attributes to extra subagent-driven code review causing timeouts or out-of-scope edits in two cases examined by Cognition. And Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment that Anthropic says had a bug affecting structured-output requests; Anthropic expects any effect to be small and to understate performance, and says the bug has been fixed.

Pricing against Opus 5.5 (per million tokens)

ItemSonnet 5.5Opus 5.5
Input tokens$2$4
Output tokens$10$20
Cache reads$0.20$0.20
Cache writes$2.50$5

Usability Analysis

The customer testimonials Anthropic published are claims by the named companies, not independent measurements, but they point in a consistent direction: fewer steps and fewer tokens. Base44 reports that across 118 real app builds Sonnet 5.5 produced apps that scored level with Opus 5, in 3.6 iterations on average versus 7.7 for Opus 5. Balyasny Asset Management reports about 121k tokens per answer on its 2,441-task finance suite, where Sonnet 5 used 497k. Lovable reports a third fewer tool calls and roughly half the shell runs. Zendesk reports tickets processed 20% faster, and Slack reports about 14% fewer output tokens on its offline Slackbot evaluations. Box reports 2.4x faster responses and 12% fewer total tokens, and Atlassian says teams can run Rovo Agents up to 30% faster.

Qualitative notes from testers include better writing than the previous generation, a knack for design, and slide decks that need minimal editing. In one Anthropic internal test, two experts judged a first-draft 10-slide operating review ready to send as is.

Practical migration points matter for existing users. Anyone running Sonnet with thinking off must switch to the new between_tools setting, which keeps up-front thinking off, before moving to Sonnet 5.5. Because preserved thinking cannot be decoupled from the account that created it, people who move conversations between accounts, including switching accounts mid-session in Claude Code, are affected. Anthropic says most developers will not notice a change. The model ID is claude-sonnet-5-5, and zero data retention is available.

Pros and Cons

Pros: unchanged per-token pricing with lower per-task cost (Anthropic's testing); a large agentic coding jump over Sonnet 5; availability on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure from day one.

Cons: Opus 5.5 still leads on most published rows and, per Anthropic, on complex open-ended work; higher-risk cybersecurity requests fall back to Sonnet 5; and the migration step for thinking-off users adds friction.

Outlook

Anthropic says Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the family "in the coming weeks," with no date given. If the Sonnet 5.5 efficiency figures hold outside Anthropic's testing, the practical question for teams becomes routing: Anthropic suggests Sonnet 5.5 pairs best with Opus 5.5 at lower effort settings, and a Creator quote in the announcement describes letting Sonnet 5.5 implement an architecture that Opus 5.5 sets. Anthropic also says cyberdefenders will soon be able to apply to an expanded Cyber Verification Program for tiered access to more advanced capabilities on Sonnet 5.5, Opus 5.5, and Claude Mythos models.

Conclusion

Sonnet 5.5 is a cost and speed release rather than a new capability ceiling, and Anthropic says as much in its alignment section. For teams running well-scoped coding, document, and support workloads on Sonnet 5, the token savings and speed gains are the reason to test it. For open-ended, judgment-heavy work, Anthropic's own guidance still points to Opus 5.5. Because nearly all figures here come from Anthropic and its customers, teams should re-run their own evaluations before switching production traffic. Rating: 4 out of 5.

Editor's Verdict

Claude Sonnet 5.5: 30% Faster, Up to 30% Cheaper per Task earns a solid recommendation within the Claude space.

The strongest case for paying attention: the per-task cost falls by up to 30% versus Sonnet 5 at unchanged per-token prices, according to Anthropic's own testing. That alone raises the bar for what readers should expect in this space. Reinforcing that, a 30%+ output speed gain over Sonnet 5 makes this Anthropic's fastest Sonnet model to date, per the company — practical value rather than just headline appeal. The broader signal worth registering is straightforward: the headline gain is efficiency rather than a higher ceiling, since Anthropic says Sonnet 5.5 does not advance the frontier of its models' capabilities while per-token prices stay at $2 and $10 per million. On the other side of the ledger, one constraint is real rather than a marketing footnote: the lead over Opus 5.5 is narrow and limited to Terminal-Bench 4.0 among the published rows, while Anthropic says Opus 5.5 remains clearly stronger on complex, open-ended work. It should factor into any serious decision. Layered on top of that, the two Artificial Analysis scores (GDPval-AA and AA-Briefcase) came from a pre-release deployment with a structured-outputs bug that Anthropic says has since been fixed — which narrows the set of teams for whom this is an obvious yes.

For Anthropic and Claude users, alignment-focused teams, and developers already invested in the Claude ecosystem, this is a serious evaluation candidate, not just a curiosity to bookmark. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.

Advertisement

Pros

  • The per-task cost falls by up to 30% versus Sonnet 5 at unchanged per-token prices, according to Anthropic's own testing.
  • A 30%+ output speed gain over Sonnet 5 makes this Anthropic's fastest Sonnet model to date, per the company.
  • Agentic coding results improve sharply, including Terminal-Bench 4.0 at 70.6% versus 10.3% for Sonnet 5 in Anthropic's table.
  • Availability on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure from launch, with zero data retention offered.

Cons

  • The lead over Opus 5.5 is narrow and limited to Terminal-Bench 4.0 among the published rows, while Anthropic says Opus 5.5 remains clearly stronger on complex, open-ended work.
  • The two Artificial Analysis scores (GDPval-AA and AA-Briefcase) came from a pre-release deployment with a structured-outputs bug that Anthropic says has since been fixed.
  • Higher-risk cybersecurity tasks visibly fall back to Sonnet 5, and some microbiology and virology requests may be flagged in error by the biology safeguards.
  • Users running Sonnet with thinking off must switch to the new between_tools setting before migrating.
Advertisement

Comments0

Key Features

1. Released September 28, 2026 as the second model in the Claude 5.5 family; model ID claude-sonnet-5-5, available on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure, with zero data retention. 2. Pricing unchanged from Sonnet 5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache reads; Anthropic says it costs up to 30% less per task and generates output 30%+ faster. 3. Scores 70.6% on Terminal-Bench 4.0 (Sonnet 5: 10.3%) and 52.1% on FrontierCode 1.1 at Xhigh effort, but trails Opus 5.5 on FrontierCode, CursorBench, GDPval-AA, Humanity's Last Exam, and OSWorld in Anthropic's table. 4. Default effort is Medium in Claude Code and the Claude apps, and High on the Claude Platform; users running Sonnet with thinking off must move to the new between_tools setting. 5. First Sonnet model to launch with cyber safeguards (higher-risk cyber tasks fall back to Sonnet 5) and with classifiers that prevent reasoning extraction; preserved thinking is expanded so thinking cannot be decoupled from the originating account.

Key Insights

  • The headline gain is efficiency rather than a higher ceiling, since Anthropic says Sonnet 5.5 does not advance the frontier of its models' capabilities while per-token prices stay at $2 and $10 per million.
  • A closer read of Anthropic's table shows Sonnet 5.5 ahead of Opus 5.5 only on Terminal-Bench 4.0, so TechCrunch's broader agentic-coding framing overstates the published numbers.
  • Token efficiency is the mechanism behind the savings, and customers such as Balyasny (about 121k tokens per answer versus 497k for Sonnet 5) report the same pattern in their own tests.
  • Effort settings change the value proposition, because Anthropic says Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score on several benchmarks for about a tenth of the cost per task.
  • The FrontierCode result at Max effort (46.2%, below the 52.1% at Xhigh) shows that more reasoning effort can lower a score when extra subagent review causes out-of-scope edits.
  • Cyber safeguards now reach the Sonnet tier, with higher-risk tasks visibly falling back to Sonnet 5, because Anthropic says the model's cyber capabilities are comparable to Opus 5's.
  • Anti-distillation measures have real workflow effects, since preserved thinking ties a conversation to its originating account, which matters for anyone switching accounts mid-session in Claude Code.
  • Anthropic's own guidance keeps Opus 5.5 as the choice for complex, open-ended work, which makes Sonnet 5.5 a routing option for well-scoped tasks rather than a replacement.

Was this review helpful?

Share

Twitter/X
Advertisement