Summary

  • On Monday, Anthropic unveiled Claude Sonnet 5.5, maintaining its pricing at $2 per million input tokens and $10 per million output tokens, which is half the cost of Opus 5.5.
  • According to Anthropic, Sonnet 5.5 achieved a score of 70.6% on Terminal-Bench 4.0, surpassing Opus 5.5's 66.4%. Independent evaluations by Artificial Analysis also favor Sonnet 5.5, scoring it at 63.6% versus 59.6% for Opus 5.5.
  • While Artificial Analysis ranks Sonnet 5.5 second to Opus 5.5, it noted that Sonnet 5.5 utilized more tokens per task than any other model in its tests.

Anthropic's latest model, Claude Sonnet 5.5, was launched on Monday as an enhancement over the previous Sonnet 5 introduced in June. The company claims this new iteration operates over 30% faster than its forerunner.

According to Anthropic, “Sonnet 5.5 excels in well-defined daily tasks, debugging, and crafting refined documents, presentations, and spreadsheets. It also demonstrates a keen eye for design.”

Myriad: How low will Nvidia go? Click to make your prediction.

The pricing structure remains unchanged at $2 per million input tokens and $10 per million output tokens. Tokens represent segments of text that the AI processes, slightly shorter than a word, and companies charge by the million. This pricing is significantly lower than that of Opus 5.5. Moreover, Sonnet 5.5 consumes nearly a third fewer tokens per task, making it more cost-effective than Sonnet 5.

This model is particularly effective in coding tasks. It scored 70.6% on Terminal-Bench 4.0—a benchmark that assesses an AI's ability to perform complex professional tasks autonomously—compared to Opus 5.5's 66.4% and Sonnet 5's 10.3%.

In simpler terms, the more affordable Sonnet 5.5 completed a greater number of tasks. An independent assessment by Artificial Analysis corroborated this, showing scores of 63.6% for Sonnet 5.5, 59.6% for Opus 5.5, and 59.1% for OpenAI's GPT-6 Astra.

Claude Sonnet 5.5 (max) shows significant progress on Terminal-Bench, ranking among the top models in both Terminal-Bench 4.0 and Terminal-Bench-Science. It achieved a score of 64%, marking a 50-point improvement over Claude Sonnet 5 (max), and slightly surpasses both Opus 5.5 and GPT-6… pic.twitter.com/7CPrhgmfxe

— Artificial Analysis (@ArtificialAnlys) September 28, 2026

Performance scores vary based on the effort setting, which allows the model to spend more time crafting answers at a higher cost. Anthropic claims that at a High effort setting, Sonnet 5.5 competes with GPT-6 Sol on FrontierCode for a fraction of the task cost.

On GDPval-AA, which evaluates real-world professional performance across 44 job categories using an Elo ranking system, Sonnet 5.5 scored 1844, closely trailing Opus 5.5's 1846. In contrast, GPT-6 Sol scored 1487.

Competitors have adjusted their pricing as well. Last week, OpenAI lowered the prices of GPT-6 Sol to $2 and $10, while its mid-tier model, GPT-5.6 Terra, is available at $2 and $12. Anthropic has not released Terra performance benchmarks.

Potential Drawbacks

Sonnet 5.5 tends to generate a large volume of text, producing approximately 193,000 tokens per test task at maximum effort, which is the highest recorded by Artificial Analysis and about 60% more than Opus 5.5. This results in a cost of $7.60 per task, about 50% higher than Sonnet 5, which contradicts Anthropic's claim of up to 30% savings.

However, Anthropic's savings are realized at lower settings. At Medium effort, the default setting in its applications, Sonnet 5.5 surpasses Sonnet 5's best coding score for less than one-tenth of the cost. Artificial Analysis suggests that High effort offers the best value. For average users, this means they can access near-flagship coding capabilities at a significantly reduced price, as long as they keep the effort level low.

It is worth noting that Anthropic's data is self-reported, and Artificial Analysis tested a pre-release version that may have contained a bug, which Anthropic believes would have had minimal impact on the scores. Anthropic maintains that Opus 5.5 remains superior for complex tasks requiring sustained judgment.

A new model, Claude Haiku 5.5, designed for high-volume, cost-sensitive applications, is expected to launch in the coming weeks.

Daily Debrief Newsletter

Stay updated with the latest news stories, original features, podcasts, videos, and more by subscribing to our daily newsletter.