Summary

  • Haiku 5.5 pairs a 1M-token context window with a fivefold price step for prompts over 100,000 tokens.
  • Because the same text yields about 30% more tokens than with Haiku 4.5, buyers need new prompt counts and workload-specific estimates; Anthropic’s average-savings figure is not a customer forecast.

Analysis

A context window is a capacity limit. It is not a promise that every token inside it costs the same. Claude Haiku 5.5, announced on 7 October, can take up to one million tokens of context and return up to 128,000 output tokens. Yet its published rates change much earlier: for prompts up to 100,000 tokens, input costs $0.10 and output $0.50 per million tokens; above that prompt length, the rates are $0.50 and $2.50. A request that crosses the line moves into a price band five times higher for both input and output. The Claude Platform price table makes that threshold explicit.

This is not a contradiction in engineering terms. A larger window makes long documents, histories and tool results available to a model in one request. A price band gives Anthropic a way to charge differently for longer prompts. But for a buyer comparing a model by its headline input rate, the distinction is material: the 1M figure describes what fits, while the 100,000-token line helps determine what it costs.

The second change complicates comparison with the previous Haiku. Anthropic says the newer tokenizer produces approximately 30% more tokens for the same text, with the exact increase depending on the content. Its migration guide tells developers to recount prompts and revisit their cost estimates rather than reuse Haiku 4.5 token measurements.

The arithmetic shows why. Suppose a text counted as 50,000 input tokens on Haiku 4.5 becomes 65,000 on Haiku 5.5. At list rates, the input-only bill would move from $0.05 to $0.0065, about 87% lower. Now take a text that counted as 80,000 tokens on the old model. A 30% increase would place it near 104,000 tokens on the new one: about $0.08 on Haiku 4.5 versus $0.052 on Haiku 5.5, roughly 35% lower. Both examples use approximate tokenizer math; they exclude output, tools, caching, discounts and any difference in task performance. They are illustrations, not customer outcomes.

Anthropic’s release says Haiku 5.5 costs around 75% less to run on average. Its footnote says the calculation includes the tokenizer change, and that 90% of requests to Haiku 4.5 were at or below 100,000 tokens. That supports a plausible average for the workload mix Anthropic measured. It does not tell a particular buyer how many of its own requests sit near the boundary, how much of its token volume comes from long prompts, or whether the traffic mix will remain the same after migration. Request counts and token-weighted spend are different distributions.

The prompt-length rate also interacts with product choices. Prompt caching has separate write and read prices, and Anthropic’s Batch API discounts input and output token prices by 50%. Repeated context may therefore behave differently from a one-off long prompt. A “million-token job” is not a single cost category: it may include cached history, tool definitions and results, documents, output and thinking tokens. Anthropic’s context-window documentation says these components count toward context use. A bill estimate that counts only the document on disk can miss the request’s actual token load.

The commercial implication is a more granular procurement question. Buyers should recount representative prompts with Haiku 5.5’s tokenizer, then group traffic below, near and above 100,000 tokens. They should price input and output separately and record cache and Batch use. If the long-prompt band drives a disproportionate share of spend, teams can test whether retrieval, summarization or selective context improves the economics without reducing the quality of the result. The savings from shortening a prompt are not free if the omitted material matters.

The useful comparison is therefore not “one million tokens versus a smaller window.” It is what the model costs across the prompt lengths a workload actually generates, after the tokenizer, cache strategy and output are counted. Haiku 5.5 makes long context available; the price ladder makes it necessary to measure how often a deployment uses it.

Sources