Claude Haiku 5.5 is up to 90% cheaper than Haiku 4.5. Your high volume automation steps just got a new default
The cheap model tier just moved, and the steps that run thousands of times a day are where you feel it first.
What happened
Anthropic released Claude Haiku 5.5, which it calls its cheapest, fastest and most capable small model. It is built for high volume work like summaries, database queries, classification and live customer support, and for running as a subagent next to Sonnet 5.5 or Opus 5.5.
Pricing is $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, rising five times above that. Simon Willison notes Haiku 4.5 cost $1 and $5, and that the new price exactly matches OpenAI's GPT-6 Luna under 100,000 tokens. Anthropic says prompts under that line made up about 90% of previous Haiku requests, and that Haiku 5.5 costs around 75% less to run on average.
There is a catch. Haiku 5.5 uses a new tokenizer, and Willison measured the same long prompt using about 1.25 times as many tokens as on Haiku 4.5. Anthropic also says the model is best for narrowly scoped tasks, with Sonnet 5.5 and Opus 5.5 still the better picks for complex agentic coding.
The most relevant customer quote for CRM builders comes from HubSpot, which said Haiku 5.5 got the best score it has seen on its CRM evaluation suite, 92.8% averaged over three runs, and was fastest with the lowest false positive rate on a stale record audit task. Anthropic also halved Sonnet 5.5 cache read prices and added monthly API credits for Max and Team subscribers.
My take
Most automations I build are not one big clever prompt. They are dozens of small steps: tag this inbound lead, summarize this call, check if this record is a duplicate, pull the deal value out of this email. Those steps run constantly and rarely need a top tier model.
That is exactly the job Haiku 5.5 is priced for. Three things I would do this week:
- List every AI step in your workflows that uses a mid or top tier model for a short, bounded task.
- Re-run a small labelled sample of real inputs through Haiku 5.5 and compare accuracy, not just cost.
- Budget by the job, not the token price. The tokenizer change means real savings will be smaller than the headline.
Keep long context jobs where they are. Above 100,000 tokens the pricing changes, and Willison points out Luna looks like a better deal there.
More posts
- LangChain built an agent that pays real merchants with Stripe's Link. The spending limit lives in code the model cannot touchOct 10, 2026
- Deno is joining Cloudflare and Deno Deploy shuts down in about six months. Check where your webhooks and scripts runOct 10, 2026
- AI coding agents added 23% more pull requests but no more finished features. Review is the bottleneck in your automations tooOct 10, 2026
