Systems
Security
Data
Python

PLEXData Engineering for people and agents that act on systems

Permissions, Limits, Evidence & eXecution.

No. 012 Latest feature

Wednesday, 30 September 2026

plexdata.online

009

Published
Author
Dorian Sotpyrc
Reading
6 minutes · News analysis
Status
Launch confirmed · vendor benchmarks unverified

News · Models

Sonnet 5.5 Turns the Mid-Tier into the Default

Anthropic’s new Sonnet posts frontier-class scores at Sonnet prices. The number that matters is not the 70.6% headline — it is the cost per task at the effort level you actually run.

References

Bar chart of Anthropic-reported Terminal-Bench 4.0 scores, per cent of tasks completed: Claude Sonnet 5.5 leads at 70.6%, Claude Opus 5.5 scores 66.4%, and Claude Sonnet 5 scores 10.3%. This is vendor benchmark data, not PLEXData testing.

Fig. 1 · Terminal-Bench 4.0Vendor data

Sonnet 5.5 leads the chart at the same token price as Sonnet 5. At default effort, Anthropic says it beats Sonnet 5’s best score for about a tenth of the cost per task.

Anthropic-reported · 28 Sep 2026
In this article · 5 sections
  1. 01What changed on 28 September
  2. 02Cost per task, not the headline
  3. 03Effort is now the knob
  4. 04Shipped vs claimed
  5. 05What to do before switching
Confirmed

Sonnet 5.5 shipped

Anthropic released Claude Sonnet 5.5 on 28 September 2026. It is available on all major platforms as claude-sonnet-5-5, priced the same as Sonnet 5: $2 per million input tokens and $10 per million output.

Vendor claim

Frontier class, a tenth of the cost

Anthropic reports 70.6% on Terminal-Bench 4.0 against Sonnet 5’s 10.3%, and says that at default effort it beats Sonnet 5’s best score for about a tenth of the cost per task. These are vendor numbers.

Open question

Your workload, your effort level

Gains depend on task mix and effort settings, and Anthropic itself says Opus 5.5 remains clearly stronger on complex, open-ended work. Independent replication is pending.

Choosing a model tier used to be a real decision: pay for the frontier model when the work is hard, drop a tier when it isn’t, and accept the gap in between. This launch is Anthropic’s attempt to close that gap — and if the numbers hold, the default for everyday agent work just moved down a tier.

On 28 September, Anthropic released Claude Sonnet 5.5, the second model in its 5.5 family after Opus 5.5 arrived on 22 September. A Haiku 5.5 for high-volume work is promised in the coming weeks.1 The pitch is blunt: a clear upgrade over Sonnet 5, 30% faster, and up to 30% cheaper per task — at exactly the same token price.

What caught my attention is not the jump itself. Model releases always jump. It is where the jump lands: the tier most teams actually afford for agents that run all day.

§ 01What changed on 28 September

The confirmed part is simple. Sonnet 5.5 is live in the Claude apps, Claude Code, and on the Claude Platform, AWS, Google Cloud and Azure.1 Pricing is unchanged from Sonnet 5, and Anthropic is publishing a system card with its evaluation method.2

API pricing, per million tokens
ItemSonnet 5.5Opus 5.5
Input$2$4
Output$10$20
Cache reads$0.20$0.20
Cache writes$2.50$5

One migration detail matters if you run agents with thinking disabled: you need to switch to the new between_tools setting, which keeps up-front thinking off, before moving to Sonnet 5.5. Anthropic’s migration guide covers it.3

§ 02The number that matters is cost per task, not the headline score

The headline number is real, and it is large. On Terminal-Bench 4.0, an agentic coding evaluation, Anthropic reports Sonnet 5.5 at 70.6% — against 10.3% for Sonnet 5 and 66.4% for Opus 5.5. On CursorBench 4.0, built from real Cursor coding sessions, it reports 55.5% against Sonnet 5’s 34.1%, within about two points of Opus 5.5.1

Anthropic-reported scores, 28 September 2026 announcement
EvaluationSonnet 5Sonnet 5.5Opus 5.5
Agentic coding · Terminal-Bench 4.010.3%70.6%66.4%
Agentic coding · CursorBench 4.034.1%55.5%57.8%
Knowledge work · GDPval-AA v2.1144918441846
Computer use · OSWorld 2.157.0%80.1%81.8%
Chart recognition · Chartography15.6%61.6%64.4%

But scores are half the story. Anthropic’s own charts plot score against cost per task at every effort level, and that is where the release bites: at Medium effort — the default in the Claude apps — it says Sonnet 5.5 beats Sonnet 5’s best score for less than a tenth of the cost per task.1 Same price list, fewer tokens per job, faster output. If that holds outside Anthropic’s harness, the cost model for routine agent work changes, not just the leaderboard.

Slack’s early testing points the same way, with the usual early-tester caution: better results than Sonnet 5 on almost all of its offline Slackbot evals, in fewer steps, with about 14% fewer output tokens — a vendor-published tester quote, not an independent result.1

§ 03Effort is now the knob, and the default moved

Sonnet 5.5 leans hard on effort levels: Low, Medium, High and beyond, trading speed and tokens against thoroughness. The defaults are Medium in Claude Code and the Claude apps, High on the Claude Platform.4

Anthropic’s own framing is that Sonnet 5.5 complements Opus 5.5 at lower effort, where it costs less per task, and can match it at higher effort at comparable cost. At High effort on FrontierCode it reports a score ten points above Sonnet 5 at the same setting — at roughly one fifteenth of the cost per task.1

My read: this makes effort a routing decision, not just a quality setting. The practical question for an agent pipeline is no longer “which model” but “which model at which effort for which job class”. Teams that pin one model and one effort for everything will pay for it twice — once in tokens, once in capability left on the table.

§ 04What shipped, what is claimed, and what nobody knows yet

Confirmed: the model exists, the model ID is live, pricing is published, the system card is public, and the safety posture is documented — the same biology safeguards as Sonnet 5, visible fallbacks to Sonnet 5 on higher-risk cybersecurity tasks, and classifiers against reasoning extraction, a first for the Sonnet line.12

Claimed: the benchmark table, the 30% speed and cost figures, the “a tenth of the cost per task” comparisons, and the collaboration anecdotes from Epic and Slack are all Anthropic-reported. Anthropic also reports that on its roughly 1,850-scenario behavioral audit, Sonnet 5.5 improves on or matches Sonnet 5 on most alignment measures, and that on containment evaluations it comes close to Opus 5.5 on how rarely it tries to escape its sandbox.1 Vendor-reported alignment is still vendor-reported.

Still open: independent Terminal-Bench and CursorBench replication; whether the cost advantage survives long-horizon, tool-heavy workloads; how the between_tools migration plays out in existing agent stacks; and where Haiku 5.5 lands the bottom tier when it ships.

§ 05What I would do before switching anything

First, re-run your own tasks at Medium effort before touching defaults. Vendor charts measure vendor-chosen tasks; your eval set is the only one that knows your workload.

Second, treat effort as part of routing. If you run an agent pipeline, this release is a good reason to log effort level alongside model and cost per task — you cannot tune a knob you do not record.

Third, keep Opus 5.5 for the work that burns. Anthropic is unusually plain that Opus remains clearly stronger on complex, open-ended work requiring sustained judgment, and benchmark scores capture only one facet of capability.1 The mid-tier became the default. It did not become the frontier.

§References

  1. 1Anthropic — Introducing Claude Sonnet 5.5anthropic.com
  2. 2Anthropic — Claude Sonnet 5.5 System Cardanthropic.com
  3. 3Anthropic — Sonnet 5.5 migration guideplatform.claude.com
  4. 4Anthropic Academy — Choosing the right effort level in Claude Codeacademy.claude.com