On the OSWorld computer-use test, the last Claude Haiku scored 15.7 percent. The one Anthropic shipped on October 7 scores 72.4 percent on the same test’s offline subset. That is not an incremental bump on a budget model; it is the cheapest, smallest tier in the Claude family suddenly doing work that was Sonnet territory a week ago. The launch copy leads with the other number, a 90 percent price cut, and that number is real too. But the price move is the part most buyers will read wrong, and the capability leap comes with a setting attached that decides whether you actually keep the savings.
Claude Haiku 5.5 is available now on the Claude Platform, Amazon Bedrock, Google Vertex and Microsoft Azure under the model ID claude-haiku-5-5. Anthropic calls it the cheapest, fastest and most capable small model it has released, aimed at high-volume, cost-sensitive jobs: summaries, compaction, database queries, classification, subagent work, and latency-sensitive tasks like live customer support and browser automation. Here is where it actually fits, and the two things the headline skips.
The price cut reached the floor, it did not break it
Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output for requests under 100,000 tokens, with cache reads at $0.01 and cache writes at $0.125. Above 100,000 tokens the rate rises to $0.50 input and $2.50 output. The old Haiku 4.5 billed $1.00 and $5.00, so for the sub-100K requests that Anthropic says make up about 90 percent of Haiku traffic, this is a 90 percent cut on both input and output. Anthropic frames the whole model as costing roughly 75 percent less to run on average.
Read the competitive context before you treat that as a price war won. The new rate does not undercut the market; it matches GPT-6 Luna’s listed rates exactly, which sat at $0.10 and $0.50 already. Against the rest of the small-model field, Anthropic’s cut moves Haiku from well above the floor to sitting on it.
| Small model (API list price) | Input / million | Output / million |
|---|---|---|
| Claude Haiku 5.5 (under 100K) | $0.10 | $0.50 |
| GPT-6 Luna | $0.10 | $0.50 |
| Google Gemini 3.5 Flash-Lite | $0.30 | $2.50 |
| Grok 4.3 | $1.25 | $2.50 |
| Claude Haiku 4.5 (previous) | $1.00 | $5.00 |
Competitor rates as reported by VentureBeat; treat list prices as a starting point, not your bill.
That distinction matters for how you plan. Anthropic did not cut Haiku’s price to win on cost; it cut the price to stop bleeding high-volume traffic to models that were already cheaper. Parity, not advantage. If you were routing bulk classification or summarization to Luna or Flash-Lite on price alone, Haiku 5.5 removes the reason to avoid it, but it does not hand you a cheaper option than the floor. The reason to choose it now is capability at the floor price, which is a different argument than “it’s the cheapest,” and it changes what you should test for.
The effort dial moved onto the cheapest model
The quieter change is the one that will shape your invoice. Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, the low / medium / high / max dial that Sonnet and Opus already carried. Medium is the default.
That dial is exactly where the launch benchmarks live, and the launch benchmarks are reported at the top of it. Anthropic’s headline agentic-coding figure, 39.2 percent on Terminal-Bench 4.0, is the score at maximum effort. At medium, the default, the same model scores roughly 20 percent. So the number that makes the budget model look capable of agentic work is the most expensive way to run the cheapest model, and it is not the number you get out of the box.
This is the same lesson the effort dial taught on the bigger models, now arriving at the tier where it bites hardest. On Sonnet 5.5 the effort setting spans close to a tenfold range in cost per task while the model underneath never changes. Put that same spread on a model whose whole reason to exist is high volume, and a Haiku instance left at max effort on a million daily calls is no longer a budget line. You can turn the $0.10 model into something that bills like a mid-tier one without changing a single routing decision, just by leaving the dial where a benchmark wants it. Budget the effort level, not the sticker.
A fair caveat on all of these figures: they are Anthropic’s own reported numbers, graded by Anthropic. Independent benchmarking from Artificial Analysis had the previous Haiku 4.5 at about 15 on its Intelligence Index, near the bottom of the small-model field, and it has not yet published an independent score for 5.5. Treat the leap as a hypothesis about your workload until you have run your own.
What the new Haiku can do that the old one could not
The capability jump is large enough to change routing decisions, and the computer-use number is the clearest signal. OSWorld measures whether a model can actually operate a desktop and browser to finish a task, and Haiku went from 15.7 percent to 72.4 percent on the offline subset. On GDPval-AA v2.1, a broad knowledge-work benchmark, it moved from 735 to 1,620. On Humanity’s Last Exam without tools it went from 10.2 to 45.9 percent. On Terminal-Bench 4.0 the prior Haiku scored 0.0; this one reaches 39.2 at max.
What that unlocks in practice is a budget model that can sit inside a browser-automation or computer-use loop and complete simple steps reliably, which previously meant paying Sonnet rates for every step. For high-volume agentic jobs where most actions are routine (navigate, extract, classify, fill a field) and only a few are hard, Haiku 5.5 can now carry the routine ones at a tenth of the cost.
The ceiling is just as important to the decision. Even at max effort, Haiku’s 39.2 on Terminal-Bench sits far below Sonnet 5.5’s 70.6 on the same test. This is not a coding model, and Anthropic says so plainly: Sonnet 5.5 and Opus 5.5 remain the choice for complex, multi-step agentic coding. Haiku got good enough to do scoped work; it did not get good enough to do hard work. Reading the OSWorld jump as “Haiku can code now” is the mistake the launch numbers invite.
Where it earns a slot, and where it does not
The practitioner move is not to route everything to the cheapest model. It is to match the model and the effort to the shape of the job, then verify the choice on your own traffic rather than Anthropic’s suite.
Haiku 5.5 earns a slot for high-volume, well-scoped, latency-sensitive work: classification, summarization, compaction, data extraction, routing, first-pass triage, and the subagent calls an orchestrator spawns by the thousand. Run it at low or medium effort and step up only when your own evals show the cheaper setting failing. Because this tier is where volume lives, the cache matters more than anywhere else: at $0.01 per million cache reads, a stable system prompt read against thousands of calls is nearly free, and a low cache-hit rate on a fixed prefix is the first thing to instrument and alert on. It does not earn a slot for complex agentic coding or hard single-shot reasoning; that is still Sonnet and Opus work. And for the narrowest bounded decisions, a “should I call this tool” or a relevance gate, even a $0.10 model is more than you need next to a purpose-built decision model that answers in one pass.
Running IT operations at a large telecom for twenty years taught me that the line item nobody set on purpose is the one that quietly grows, and the cheapest unit multiplied by the largest volume is exactly that line item. A subagent fleet left at the default effort is a fleet-wide cost decision made by a default, not by you. Pin the effort level deliberately, meter cost per completed task rather than price per token, and test the capability leap on fifty to two hundred of your own tasks at two or three effort levels before you trust a launch chart. The 90 percent cut is genuine, but it brings Haiku to the floor the competition already set. The thing worth paying attention to is not that the price fell; it is that the dial controlling your savings now sits on the model you will run the most.
