Sonnet 5.5 narrows the gap with Opus on coding and agentic work
Anthropic is expanding the Claude 5.5 family with Sonnet 5.5, a model designed to run faster and at a lower cost per task than Sonnet 5. Its largest gains focus on agentic coding, Computer Use, visual understanding, and professional work, alongside new safeguards for cyber capabilities.
30% faster at the same token pricing
The Claude 5.5 family now has a second model. Sonnet 5.5 retains Sonnet 5 pricing at $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache-read tokens.
Anthropic says the model typically requires fewer tokens to complete the same work. In its testing, that can reduce cost per task by up to 30%. Output generation is also reported to be more than 30% faster than Sonnet 5.
Its positioning remains distinct from Opus 5.5. Anthropic reserves its higher-tier model for complex work requiring sustained judgment, while Sonnet targets well-scoped everyday tasks, bug fixing, and document, presentation, and spreadsheet creation.
Haiku 5.5 is expected to complete the family in the coming weeks for high-volume and cost-sensitive applications. A major jump on Terminal-Bench 4.0
The most pronounced change appears in agentic coding. Anthropic reports a 70.6% score for Sonnet 5.5 on Terminal-Bench 4.0, compared with 10.3% for Sonnet 5.
On FrontierCode 1.1, Sonnet 5.5 reaches 46.2% at its reported best effort level, compared with 42.4% for Sonnet 5 and 49.3% for Opus 5.5. CursorBench 4.0 places Sonnet 5.5 at 55.5%, close to Opus 5.5 at 57.8%.
Anthropic also observed a change in how the model uses tools. In head-to-head testing, Sonnet 5.5 batched independent tool calls more frequently instead of executing them sequentially, reducing both the number of steps and token usage. Computer Use moves closer to Opus
The gap also narrows on tasks outside coding. Sonnet 5.5 reaches 80.1% on OSWorld 2.1 for Computer Use, compared with 57% for Sonnet 5 and 81.8% for Opus 5.5 in Anthropic's reported results.
On Chartography, which evaluates chart understanding, Sonnet scores 61.6%, compared with 15.6% for its predecessor and 64.4% for Opus.
The model also records 1,844 points on GDPval-AA v2.1, an evaluation based on professional tasks across 44 occupations and nine industries. Opus 5.5 scores 1,846 on the same evaluation.
Anthropic cautions that proximity on individual benchmarks does not make the models interchangeable. Its internal and external testing continues to place Opus ahead on complex, open-ended work requiring sustained judgment. Effort becomes a cost lever
Sonnet 5.5 retains multiple effort settings to control how much time the model spends reasoning. Claude Code and the Claude apps default to Medium, while Claude Platform defaults to High.
Anthropic reports that on several benchmarks, Sonnet 5.5 running at Low or Medium effort exceeds Sonnet 5's best score for roughly one tenth of the cost per task.
That relationship changes as effort increases. Sonnet 5.5 can approach Opus 5.5 performance at higher settings, but its execution cost can move closer as well. Opus-style safeguards reach Sonnet
Cyber capabilities have improved enough for Anthropic to apply safeguards similar to those used with Opus 5.5 to a Sonnet model for the first time.
Routine software development remains available, while some higher-risk cybersecurity requests can fall back to Sonnet 5. Anthropic also plans to expand its Cyber Verification Program with tiered access to capabilities across Sonnet 5.5, Opus 5.5, and Claude Mythos.
Sonnet 5.5 also introduces classifiers intended to limit large-scale distillation attacks designed to extract model capabilities. Biology safeguards remain the same as those used for Sonnet 5.
Claude Sonnet 5.5 is available across Claude platforms as well as Amazon Web Services, Google Cloud, and Microsoft Azure. Its model ID on Claude Platform is `claude-sonnet-5-5`.