Haiku 5.5 moves AI pricing battle to the agent workhorse layer


Model routing
The practice of sending different parts of a workflow to different models based on cost, latency, difficulty and risk.
Subagent
A smaller or specialized model call that handles a bounded task inside a larger agent workflow, such as classification, extraction or validation.
Cache read
A discounted reuse of previously stored prompt context, often used to lower the cost of repeated agent calls that share the same background information.
Token threshold
A pricing boundary based on prompt size; for Haiku 5.5, prompts above 100,000 tokens move to a higher per-token rate.
Small-model push
Anthropic positioned Claude Haiku 5.5 as its fastest and cheapest small model for high-volume tasks such as summaries, classification, database queries, browser use and subagent work.
Lower pricing
Haiku 5.5 starts at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens.
Routing layer
The launch points to a model-routing architecture where larger models plan and smaller models execute repetitive agent subtasks at scale.
Anthropic released Claude Haiku 5.5 on October 7, positioning it as its fastest and lowest-cost small model for repetitive, high-volume work, including summaries, context compaction, database queries, classification, browser use and subagent tasks.1 The launch is a pricing move, but its larger implication is architectural: AI product teams are being pushed toward model-routing systems in which premium models plan and decide, while cheaper small models execute the narrow tasks that make agents expensive in production.
Haiku 5.5 starts at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, with cache reads at $0.01 per million tokens in that tier.1 Above 100,000 tokens, the model moves to $0.50 per million input tokens and $2.50 per million output tokens.2 Anthropic says Haiku 5.5 costs about 75% less to run than Haiku 4.5 on average, while the largest list-price cut — 90% — applies to requests below the 100,000-token threshold.1
For AI product builders, the headline is not just cheaper tokens. It is that small models are becoming the default execution layer for agentic systems. Anthropic’s launch materials describe Haiku 5.5 as a fit for “quick and repetitive workloads” and as a subagent paired with Sonnet 5.5 or Opus 5.5 for coding work.1 The platform documentation reinforces that routing role, listing classification, extraction, routing and subagent tasks among the model’s intended uses.2
The key product-design question raised by Haiku 5.5 is where to draw the line between expensive reasoning and cheap execution. A flagship model can still handle hard planning, ambiguous judgment and final synthesis. But many agent workflows are dominated by smaller operations: classify this ticket, summarize this page, extract a field from this filing, check this tool result, rewrite this block or choose the next handler.
Those tasks are increasingly measurable, repeatable and routable. That makes them candidates for a small model, provided reliability is high enough and the savings justify the orchestration overhead. AWS framed Haiku 5.5 this way in its Bedrock launch post, describing it as a fast subagent layer for routing, classification, summarization, document rewriting and small code changes, while Opus 5.5 handles planning and harder judgment calls.4
That division of labor is likely to become a standard agent pattern: a larger model decomposes work, a router assigns subtasks, and cheaper models run in parallel on bounded jobs. In that setup, a small model’s value is not just its standalone intelligence. It is its cost per successfully completed subtask, including retries, latency, tokenization effects, cache use and failure handling.
Haiku 5.5’s pricing is explicitly shaped around short and mid-sized requests. Anthropic says prompts up to 100,000 tokens make up about 90% of requests to the previous Haiku model, and it gives that tier the deepest discount.1 That matters because many agent systems generate large numbers of short calls rather than a few monolithic prompts.
VentureBeat noted that the lower tier matches OpenAI’s GPT-6 Luna list price at $0.10 per million input tokens and $0.50 per million output tokens, while emphasizing that Haiku 5.5’s higher tier begins after 100,000 tokens.5 The Frontier made the same routing-relevant point: Haiku 5.5 looks especially competitive until prompts pass the 100,000-token boundary, after which long prompts cost five times the short-rate price.8
That gives builders a practical incentive to break agent systems into smaller calls where possible. Instead of stuffing an entire corpus, conversation history or tool trace into one request, teams may get better economics by caching stable context, extracting only relevant spans and routing discrete decisions to the cheapest capable model. The model-routing layer becomes a cost-control surface, not just a quality-control surface.
The caveat is that token price alone is not the bill. Anthropic’s documentation says Haiku 5.5 uses a newer tokenizer and that the same text counts as approximately 30% more tokens than on Haiku 4.5.2 Teams migrating to the new model should remeasure real prompts rather than applying headline discounts directly to prior invoices.
Haiku 5.5 is not being marketed only as a text utility model. Anthropic says it is suited to browser use, live customer support and subagent work, and the AWS post says it can serve as a computer-use subagent for repetitive browser and desktop tasks.14 That positioning reflects a broader change: small models are being asked to do operational work inside software loops, not just generate short text completions.
The platform specs support that role. Haiku 5.5 has a 1 million-token context window, up to 128,000 output tokens, adaptive thinking, a medium default effort setting and the model ID claude-haiku-5-5 across the Claude API and major cloud platforms.2 It is also available on Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS, according to Anthropic’s documentation.2
Those capabilities give builders more room to design tiered systems. A product might use Haiku 5.5 for intent routing, retrieval triage, draft summaries and validation checks, while reserving Sonnet or Opus for high-risk decisions, complex coding, long-horizon planning or final customer-facing outputs. The goal is not to replace flagship models. It is to reduce how often they are invoked.
Anthropic’s benchmark table shows Haiku 5.5 ahead of Haiku 4.5 across listed evaluations and competitive with GPT-6 Luna in the tests Anthropic published, while Sonnet 5.5 remains stronger on many harder tasks.1 SiliconANGLE highlighted the same pattern, noting Haiku 5.5’s stronger reported scores than Luna on shared Anthropic benchmarks and Anthropic’s continued recommendation of Sonnet 5.5 and Opus 5.5 for complex agentic coding.6
For product teams, the main lesson is to test at the task level. A general benchmark score does not show whether a model can classify a company’s support tickets, extract revenue figures from filings or safely click through a browser workflow with the required accuracy. The Frontier also flagged an important evaluation caveat: headline benchmark results may use higher effort settings than the default, so teams should test the effort level they plan to run in production.8
That makes routing evaluation more granular. Builders need per-task acceptance thresholds, fallback rules and cost curves. A small model can be the right first pass if it succeeds often enough and fails safely. If failures require expensive retries or human review, the apparent token savings can disappear.
Anthropic also cut Claude Sonnet 5.5 cache-read pricing by half, from $0.20 to $0.10 per million tokens, saying the change should make Sonnet 5.5 about 20% cheaper on most agentic work because cached context can make up a large share of token consumption.1 SiliconANGLE and BitsMinds both treated the Sonnet cache reduction as part of the same cost-efficiency push around agent workloads.69
That matters because routing is not only about choosing a model. It is also about deciding what context to reuse, what to compress, when to call a larger model and how to amortize repeated work. Cache pricing, prompt-length thresholds and small-model rates are converging into one optimization problem.
Anthropic’s release notes also say Max and Team plans are receiving monthly API credits for apps and agents on the Claude Platform, further steering customers toward production-style API workflows rather than only chat usage.3
The immediate takeaway is to audit agent traces. Teams should identify which calls involve judgment and which are repetitive transformations, lookups or routing decisions. The second group is where Haiku 5.5-style pricing can matter most.
A practical migration plan should include four checks. First, re-tokenize real prompts under Haiku 5.5 because token counts may rise versus Haiku 4.5.2 Second, measure how many calls cross the 100,000-token threshold, since that boundary changes the economics sharply.18 Third, run local evaluations by task and effort setting, not just by model name. Fourth, design fallback paths so a larger model handles ambiguity, low-confidence results and high-impact decisions.
The broader signal is clear: the next AI pricing battle is moving below the flagship tier. As agents become more common, the workhorse layer — the small models that classify, summarize, route, browse, validate and extract thousands of times per workflow — will increasingly determine gross margins and user latency. Haiku 5.5 is Anthropic’s latest entry in that fight, and it gives builders another reason to treat model routing as core product infrastructure rather than an implementation detail.
Comments