❯ Anthropic releases Claude Haiku 5.5, priced 90% below its predecessor for requests up to 100,000 tokens
A big cut for the small modelAnthropic released Claude Haiku 5.5 on October 7, the smallest and fastest tier in its lineup. For requests up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens; above that, $0.50 and $2.50. The previous Haiku 4.5 cost $1 and $5. The company says overall running costs are about 75% lower.
The first Haiku with an effort dialHaiku 5.5 is the first Haiku to offer effort settings, with five levels from low to max, letting developers trade cost against accuracy. Anthropic recommends it for high-volume work such as summaries, classification, database queries, live customer support, browser use and subagents, and says plainly that Sonnet 5.5 and Opus 5.5 remain better for complex agentic coding. The model is live on the Claude Platform and on AWS, Google Cloud and Microsoft Azure.
Cheap, with conditionsDeveloper Simon Willison notes in his write-up that the new price matches OpenAI’s GPT-6 Luna, but only up to 100,000 tokens, beyond which Luna is a much better deal. The new model’s tokenizer also differs, so the same prompt uses roughly 1.25 times as many tokens as before. The same day, Anthropic halved the cache-read price for Sonnet 5.5 and began giving Max and Team subscribers monthly API credits.
Who switchesFor companies running large volumes of short requests every day, such as support triage or document Q&A, the cut changes the bill directly. Among customers Anthropic cites, Asana reports latency down more than 30% and Box an 11-point score gain at half the latency. For long-document workloads whose prompts regularly exceed 100,000 tokens, the savings are much smaller.
▮ SIGNALTwo leading labs have now priced their small models at the same level, and competition at the low end has moved from who is cheaper to where the conditions on cheapness are written.