Claude 5.5 Haiku drops the entry price of the Haiku range by 90%

Claude Haiku 5.5 cuts its prices by 90% below the 100,000-token threshold. What this reduction changes for the deployment of multi-agent architectures.

Anthropic is completing its Claude 5.5 lineup with Claude Haiku 5.5, introduced on October 7, 2026 as its fastest, cheapest and most capable small model to date. Haiku remains focused on high-volume workloads where latency and cost matter more than maximum intelligence, including summarization, context compaction, classification, database queries, live customer support and browser use.

The biggest change from Haiku 4.5 is pricing. For prompts up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, compared with $1 and $5 respectively for Haiku 4.5. Cache reads drop to $0.01 per million tokens, while cache writes cost $0.125. Beyond 100,000 prompt tokens, input and output pricing increases to $0.50 and $2.50.

Anthropic therefore estimates that Haiku 5.5 costs roughly 75% less to run on average than Haiku 4.5. For requests below the 100,000-token threshold, which accounted for around 90% of previous Haiku usage, the headline token pricing is 90% lower. The company's calculation also accounts for an updated tokenizer, which can slightly change the number of tokens required to complete the same task.

The lower price does not simply come from positioning Haiku as a less capable tier. In evaluations published alongside the model, Haiku 5.5 reaches 72.4% on the offline subset of OSWorld 2.1 for computer use, compared with 15.7% for Haiku 4.5 and 48.9% for GPT-6 Luna. It also reaches 45.9% without tools on Humanity's Last Exam and 39.2% on Terminal-Bench 4.0. These figures come from Anthropic's own evaluation setup, so direct comparisons across providers should still be treated with some caution.

Haiku 5.5 is also the first Haiku-class model to include an adjustable effort setting. Developers can choose between shorter, cheaper execution or allow the model to spend more resources on a harder problem. Anthropic also positions it as a subagent alongside Claude Sonnet 5.5 and Opus 5.5, allowing repetitive or narrowly scoped coding operations to be delegated without invoking the larger models for every step.

Speed is another part of the pitch. Haiku 5.5 is Anthropic's fastest model at standard model speeds, although Opus models can still run faster in Fast Mode. In early testing cited by Anthropic, Asana reported more than a 30% reduction in task-completion latency for parts of its AI Teammates product and up to 2.5x faster inference per agent turn. Those figures are customer results included in the launch material rather than independent benchmarks.

Anthropic is not presenting Haiku as a replacement for its larger models. Sonnet 5.5 and Opus 5.5 remain substantially stronger on complex agentic coding workloads such as Terminal-Bench. Instead, Haiku is designed for the large number of smaller operations that increasingly sit inside multi-agent systems: compressing context, classifying information, querying data, summarizing documents or completing a subtask before passing the result back to a larger model.

The launch also brings a pricing change elsewhere in the lineup. Anthropic is cutting Sonnet 5.5 cache-read pricing in half, from $0.20 to $0.10 per million tokens. Because cached context represents a significant share of token usage in agentic workflows, the company estimates that this makes Sonnet 5.5 around 20% cheaper for most agent workloads.

Anthropic is also introducing monthly Claude Platform API credits for selected subscriptions. Max 5x users are set to receive $100 per month, Max 20x subscribers $200, while Team plans can receive up to $500 pooled across users. The credits can be spent across Anthropic's model lineup and are intended to make it easier for subscribers to experiment with applications and agents using the API.

On the developer side, Anthropic is updating its Python and TypeScript SDKs with beta support for computer use and browser use. Haiku 5.5 is positioned as particularly well suited to these workloads because of its combination of latency, capability and price. The model is available under the `claude-haiku-5-5` identifier through the Claude Platform as well as Amazon Web Services, Google Cloud and Microsoft Azure.

The release also comes with updated safeguards. Anthropic says Haiku 5.5 shows fewer instances of misaligned behavior and lower willingness to cooperate with misuse than Haiku 4.5. Its cybersecurity protections are somewhat less restrictive than Sonnet 5.5's to support a broader range of defensive work, while still blocking activities such as penetration testing that may carry greater misuse potential. Eligible security organizations can apply for broader access through Anthropic's Cyber Verification Program.

Haiku 5.5 ultimately reflects a broader shift in how small models are being positioned. They are no longer simply cheaper, weaker versions of flagship systems. As agent architectures increasingly distribute work across several models, Anthropic is turning Haiku into a specialized execution layer: capable enough to handle a large share of routine work, fast enough for real-time interactions and inexpensive enough to be called repeatedly inside a single workflow.