Skip to content
Back to blog

Opus 5.5 and Sonnet 5.5 — judgment on one side, throughput on the other

Opus 5.5 (September 22) at Fable 5.1 level, 40% cheaper to run than Opus 5. Sonnet 5.5 (September 28) at Sonnet 5 prices, 30% faster, level with Opus on scoped work.

10 min read
  • Claude
  • Opus
  • Sonnet
  • Anthropic
  • Agents

In six days Anthropic shipped the first two models in the Claude 5.5 family. Opus 5.5 on September 22, Sonnet 5.5 on September 28. Opus is the first release since the call to pace the frontier: the lab puts it at Fable 5.1 level on most work, 40% cheaper to serve than Opus 5. Sonnet is the complement — scoped tasks, bugs, documents, slides, spreadsheets, and a sharp eye for design. Haiku 5.5 is named for the coming weeks, with no public price or model id. Benchmark margins at this level are a fragile guide: the Opus post says the felt gap with Fable is narrower than the scores. I read the tables as a vendor ceiling, effort and safeguards included.

Price and speed

Per million tokens. Opus 5.5: input $4, output $20, cache reads $0.20, cache writes $5. Opus 5 was $5 / $25 / $0.50 / $6.25. Opus fast mode (up to 2.5×, Claude Code and the Platform) is $8 / $40. Sonnet 5.5: input $2, output $10, cache reads $0.20, cache writes $2.50. Input, output, and cache reads match the Sonnet 5 rate card.

  • On a typical workload, Opus 5.5 costs 40% less than Opus 5 (cheaper tokens, and fewer tokens per task). Sonnet 5.5 costs up to 30% less per task than Sonnet 5, mostly because it writes less.
  • Both generate more than 30% faster than their predecessor. Sonnet 5.5 is the fastest Sonnet to date.
  • Cache reads, which dominate a coding agent’s bill, are the same price: $0.20. The remaining gap is input, output, and cache writes ($2.50 versus $5).
  • With Opus, Anthropic raises five-hour limits on Pro, Max, Team, and seat-based Enterprise, plus a rate-limit reset you can save for later.

Coding

On several evals, Sonnet 5.5 at max effort sits next to Opus. Anthropic adds that, in its own testing and with external testers, Opus stays clearly stronger once the work is open-ended and needs judgment held over time. The default in Claude Code and the apps is medium; on the Platform it is high. Sonnet complements Opus best at lower effort. At higher settings, scores and cost move together.

  • Terminal-Bench 4.0: Sonnet 5.5 70.6% · Opus 5.5 xhigh 66.4% · Astra high 57.9% · Fable 55.8% · Opus 5 52.3% · Sonnet 5 10.3%. The New Stack reports that the benchmark maintainers saw Sonnet 5 hit timeouts and token limits, which helps read the 10.3%.
  • FrontierCode v1.1 (would the diff merge, out-of-scope penalized): Opus 54.4% max / 54.6% medium, above Astra’s published best (53.3%) for about a fifth of the cost. Sonnet 52.1% xhigh / 46.2% max · GPT-6 Sol 49.3% · Sonnet 5 42.4%. Max landing under xhigh is documented: at max, Sonnet more often runs Claude Code’s review skill across subagents. In two cases Cognition examined, a timeout or edits past the task’s scope.
  • CursorBench 4.0: Opus 57.8% max / 52.5% medium · Sonnet 55.5% · Fable max 51.8% · Opus 5 max 46.6% · Sonnet 5 34.1%.
  • Announced cost: at high effort, Sonnet scores 10 points above Sonnet 5 on FrontierCode for about a fifteenth of the cost. At low effort on CursorBench, and at medium effort on Terminal-Bench, it beats Sonnet 5’s best for less than a tenth of the cost.

Magnitudes cited by Anthropic, not measurements of mine. A 680,000-line migration in under a day with Opus 5.5. A 200,000-line audit in under three hours, where Opus 5 took over 20 hours and 2.5× the tokens. HAProxy rewritten from C to Rust: nearly the full regression suite for both Opus 5.5 and Fable, Opus in 9.5 hours versus 12, 51% cheaper. At Lovable, across 118 builds, Sonnet 5.5 reached Opus 5’s level in 3.6 iterations on average, where Opus 5 took 7.7.

Knowledge work

  • GDPval-AA v2.1 (44 occupations): Opus 1846 Elo · Sonnet 1844 · Fable 1735 · Opus 5 1708 · Sonnet 5 1449.
  • AA-Briefcase v1.1: Opus 1822 · Sonnet 1811 · Sonnet 5 1359.
  • Humanity’s Last Exam, with tools: Opus 67.7% · Fable 65.6% · Sonnet 64.5% · Sonnet 5 54.9%.
  • OSWorld, partial: Opus 81.8% (2.0 in its own post, 2.1 in the Sonnet table) · Sonnet 80.1% on 2.1 · Sonnet 5 57.0% · Opus 5 74.0% on 2.0.
  • Chartography with no tools: Opus 64.4% · Sonnet 61.6% · Sonnet 5 15.6%. With tools, Opus reports 89.0%.
  • Where Opus 5.5 does not lead: Terminal-Bench-Science 0.1, Astra 64.6% versus Opus 58.7% (no Sonnet 5.5 figure). AutomationBench, Astra 41.4% versus Opus 40.0% — and that 40% is a floor, because a safeguard intervention counts as a failure.

Sonnet 5.5 is also the first Sonnet to beat Pokémon Red from screenshots alone. Both 5.5 models write more clearly than the previous generation: the important part first, less jargon. On a long session that weighs as much as two points of CursorBench.

Safeguards

Opus 5.5 posts the best score to date on Anthropic’s behavioral audit. Less likely to take hard-to-reverse actions, more resistant than Opus 5 to prompt injection — tying Fable at Gray Swan for the lowest success rate. Evaluated before release by Frontier Design and METR. Biology and cyber capabilities comparable to Mythos 5.1, so the safeguards match Fable’s. Life Sciences Verification is open; Cyber Verification expands in the coming weeks. Opus benches run with those safeguards on: when they intervene, cyber tasks are finished by Opus 4.8, and biology plus frontier-LLM development by Opus 5. Anthropic says that likely lowers the scores.

Sonnet 5.5 does not push the lab’s capability frontier. Across ~1,850 scenarios it matches or beats Sonnet 5 on most alignment measures. It is the Anthropic model least likely to probe its container, and it comes close to Opus in how rarely it tries to leave the sandbox. Cyber at Opus 5’s level: the first Sonnet shipped with these safeguards, and high-risk requests fall back to Sonnet 5. Biology safeguards stay the Sonnet 5 set. It is also the first Sonnet with distillation-prevention classifiers and expanded preserved thinking: reasoning stays tied to the account that produced it.

How I split them

I have not yet replayed my own loops (review, refactors, MCP) on both model ids. The split is a reading of the announcements.

  • Sonnet 5.5 at low or medium: scoped bugs, review, docs, slides, UI. That is where $2 / $10 and the speed show up.
  • Opus 5.5: migrations, audits, overnight multi-repo runs, judgment when the spec is fuzzy.
  • If Sonnet is already at xhigh or max and the bill meets Opus, I switch. The 70.6% on Terminal-Bench does not replace that switch.
  • API: claude-sonnet-5-5. If thinking was off, switch to between_tools before migrating. Both are on AWS, Google Cloud, and Azure, with zero data retention.

Takeaway

A flagship that got cheaper to run, and an everyday model that caught up on the scores. Opus 5.5 does Fable’s work on most tasks, faster, 40% cheaper than Opus 5, ahead on agentic coding, computer use, and GDPval — Astra still leads AutomationBench and Terminal-Bench-Science. Sonnet 5.5 keeps the Sonnet 5 rate card, writes less, answers 30% faster, and sits next to Opus on GDPval, CursorBench, and OSWorld. Long-horizon judgment stays with Opus. Haiku 5.5 will show whether the bottom of the lineup follows.

  • Opus 5.5: https://www.anthropic.com/claude-opus-5-5
  • Sonnet 5.5: https://www.anthropic.com/claude-sonnet-5-5