Claude Haiku 5.5 Drops to $0.10 per Million Tokens, but Prompts Over 100K Cost Five Times More

News Summary
Anthropic released Claude Haiku 5.5 on October 7, 2026, at about 11:08 a.m. Pacific Time, positioning it as the smallest and fastest model in its Claude 5.5 family and cutting the entry price for small-model API access to $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. That is roughly 90 percent below the $1 and $5 rates of Claude Haiku 4.5. The model is available under the identifier claude-haiku-5-5 on Anthropic's own platform and through Amazon Web Services, Google Cloud and Microsoft Azure.
What Anthropic Announced
Haiku 5.5 arrives about nine days after Anthropic first previewed the model alongside Sonnet 5.5 in late September 2026. According to coverage from VentureBeat, MarkTechPost, DataCamp and Unite.AI, the release centers on three ideas: a much lower base price, a large jump in measured capability over Haiku 4.5, and a set of agent-oriented features such as adjustable effort levels and beta computer-use tools. Secondary coverage also reports a context window of up to 1 million tokens and output of up to 128,000 tokens, although VentureBeat's article does not state a context size, so developers should confirm the limits in Anthropic's documentation.
Anthropic also says that, once changes in token consumption and typical request sizes are factored in, real workloads should cost about 75 percent less than on Haiku 4.5, a smaller figure than the headline 90 percent price cut.
Pricing Tiers Explained
Haiku 5.5 uses two price tiers keyed to prompt length, expressed per million tokens:
- Prompts up to 100,000 tokens: $0.10 input, $0.50 output, $0.01 for cache reads and $0.125 for five-minute cache writes.
- Prompts above 100,000 tokens: $0.50 input, $2.50 output, $0.05 for cache reads and $0.625 for cache writes.
- Claude Haiku 4.5 for comparison: $1.00 input, $5.00 output, $0.10 cache reads and $1.25 cache writes.
The Batch API reportedly takes a further 50 percent off. Several analysts have pointed out the "100K token catch": a request that crosses the threshold is billed at five times the lower rate, so teams that routinely send long documents or large codebases may see far smaller savings than the headline suggests. VentureBeat notes that its article does not say how requests of exactly 100,000 tokens are billed. Anthropic also lowered Sonnet 5.5 cache-read pricing from $0.20 to $0.10 per million tokens.
Benchmark Results
The figures below are vendor-reported and have not been independently reproduced.
- GDPval-AA v2.1, a knowledge-work evaluation: Haiku 5.5 scores 1,620, against 735 for Haiku 4.5, 1,437 for OpenAI's GPT-6 Luna and 1,840 for Sonnet 5.5.
- AA-Briefcase v1.1: 1,578 for Haiku 5.5, 614 for Haiku 4.5, 1,336 for GPT-6 Luna and 1,824 for Sonnet 5.5.
- OSWorld 2.1 (offline subset), a computer-operation test: 72.4 percent for Haiku 5.5, 15.7 percent for Haiku 4.5, 48.9 percent for GPT-6 Luna and 83.9 percent for Sonnet 5.5.
- Terminal-Bench 4.0: 39.2 percent for Haiku 5.5 at maximum effort, 0.0 percent for Haiku 4.5, 16.4 percent for GPT-6 Luna and 70.6 percent for Sonnet 5.5. At medium effort, Haiku 5.5 scores about 20 percent.
- FrontierCode 1.1: 46.4 percent for Haiku 5.5, 42.4 percent for GPT-6 Luna and 52.1 percent for Sonnet 5.5 at high effort.
An independent aggregator, BenchLM, gives Haiku 5.5 an overall score of 66.3 out of 100, ranking it 28th among 216 tracked models. That score uses a different methodology and is not directly comparable with Anthropic's tables.
Features and Safeguards
Haiku 5.5 supports adjustable effort levels, with medium as the default, so developers can trade speed and cost against depth of reasoning. Anthropic describes it as its fastest model at standard speeds, though no tokens-per-second figure was published. Beta browser and computer-operation capabilities are available through the Python and TypeScript SDKs.
The model ships with tighter cybersecurity restrictions than Haiku 4.5. Penetration-testing requests are blocked under standard safeguards, while broader access is offered to vetted users through verification programs. Anthropic also noted monthly API credits bundled with subscriptions: $100 for Max 5x, $200 for Max 20x, and up to $500 shared across a Team plan.
Alex Wang of applied AI at the financial-technology company Rogo described the model's role in enterprise workflows this way: "It's accurate enough that we'd trust it there and fast and cheap enough that we can run it a lot."
How It Compares on Price
VentureBeat's comparison table lists the following input and output rates per million tokens. It compares prices only, not capabilities.
- Claude Haiku 5.5: $0.10 and $0.50 for prompts under 100K tokens.
- GPT-6 Luna: $0.10 and $0.50, with a surcharge above 272K input tokens.
- Gemini 3.5 Flash-Lite: $0.30 and $2.50.
- Gemini 3.8 Flash: $0.75 and $3.75, promotional pricing through December 31, 2026.
- Grok 4.3: $1.25 and $2.50 under 200K prompt tokens.
- Grok 4.7: $2.00 and $6.00 under 200K prompt tokens.
- GPT-6.1 Sol and Claude Sonnet 5.5: $2.00 and $10.00.
Why It Matters for Developers and Learners
Small models handle the high-volume, repetitive parts of AI applications: classifying support tickets, extracting fields from documents, routing requests, and running sub-tasks inside larger agent systems. At $0.10 per million input tokens, a million short requests of a few hundred tokens each can cost only a few dollars, which lowers the barrier for students, startups and researchers experimenting with AI worldwide.
The pricing structure also shows a practical engineering lesson. Cost depends not just on the per-token rate but on prompt length, caching and batching. Developers can keep prompts under the 100K threshold, reuse cached context at a tenth of the input price, and use batch processing for non-urgent jobs to get the most from the new model.
Caveats
Most figures in this report come from coverage published October 7 and 8, 2026, rather than from Anthropic's pricing page, and benchmark numbers are supplied by the vendor. Prices, limits and the context window should be verified against Anthropic's official documentation before any budgeting or migration decision.
This article was compiled by the AIBARS editorial team with AI assistance. AI can make mistakes, so please check the original source for anything important. Spotted an error? Email [email protected].