Picture the engineer who runs a support assistant for a mid-sized online retailer. Every incoming ticket gets classified, summarised and routed by a small language model before a human sees it. On October 7 she opens Anthropic's announcement of Claude Haiku 5.5 and does the only calculation that matters to her: what does my monthly bill become, and does the quality hold?
The answer to the first question is unusually clear. For prompts up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. Haiku 4.5 cost $1.00 and $5.00. Cache reads drop from $0.10 to $0.01. Anthropic puts the average saving at about 75%; its own footnote says roughly 90% for short requests and 50% for long ones.
The fine print on that 75%
Two details explain why the headline number is not simply "ten times cheaper." First, pricing has two tiers. Prompts above 100,000 tokens cost five times more: $0.50 input, $2.50 output. Second, Haiku 5.5 uses the newer tokenizer introduced with Opus 4.7, which splits the same text into roughly 30% more tokens, according to Anthropic's migration notes as reported by The Daily Brief. An independent estimate that bakes in the extra tokens lands at around 87% savings for short requests and 35% for long ones. The direction is the same either way: short, frequent calls win big; long-document workloads win less.
The same day, Anthropic halved Sonnet 5.5's cache-read price from $0.20 to $0.10, which it says makes most agentic workloads about 20% cheaper. The Batch API halves Haiku's prices again, per the same third-party report.
What the benchmarks say
Anthropic's own table lists Haiku 5.5, Haiku 4.5, GPT-6 Luna and Sonnet 5.5 in that order:
- Terminal-Bench 4.0 (agentic tasks in a terminal): 39.2% / 0.0% / 16.4% / 70.6%
- OSWorld 2.1 (computer use, offline subset): 72.4% / 15.7% / 48.9% / 83.9%
- GDPval-AA v2.1 (knowledge-work tasks, Elo): 1620 / 735 / 1437 / 1840
- Humanity's Last Exam, no tools: 45.9% / 10.2% / n.a. / 56.9%
These are vendor numbers; no independent evaluation had been published at the time of writing. Still, two things stand out. The gap between the two Haiku generations is enormous, with Haiku 4.5 scoring zero on the terminal benchmark where its successor reaches 39%. And the gap to Sonnet 5.5 remains wide on coding and long agentic work, which Anthropic acknowledges by recommending Sonnet 5.5 and Opus 5.5 for complex coding and positioning Haiku for summarisation, context compaction, subagents, live customer support and browser use.
Haiku 5.5 is the first Haiku with adjustable effort levels (Low to Max). Anthropic calls it its fastest model at standard speed. Customer quotes on the launch page, which are vendor-selected, include Box reporting roughly half the latency of Haiku 4.5 and Asana reporting a latency drop of more than 30% on task completions. The model is available on the Claude API, AWS, Google Cloud and Microsoft Azure. The announcement does not state the context window or which claude.ai plans include the model.
Back to the support desk
Our engineer's assistant handles about 20 million input and 4 million output tokens a month. On Haiku 4.5 that was $40. On Haiku 5.5, even after adding 30% for the new tokenizer, it comes to roughly $5. At that level the model's price stops being a decision factor at all; accuracy becomes the only one.
In our own integration work the pattern we see most often is teams running a mid-tier model for tasks a small one could handle, simply because the small one used to be unreliable. Haiku 5.5 is the first small model in this family where that habit is worth re-examining. Our suggestion: move classification, summarisation and routing to Haiku 5.5 and measure accuracy against your own test set before and after. Keep complex, multi-step work on Sonnet 5.5, and let Haiku run the subtasks beside it. The benchmark table is a guide; your own data is the test.
Sources: Anthropic, Claude Haiku 5.5 announcement, The Daily Brief, pricing and migration analysis




