Released October 7, 2026 — vendor figures verified October 8, 2026
Claude Haiku 5.5 Benchmarks & Pricing
Anthropic’s cheapest and fastest tier, rebuilt: $0.10/$0.50 per million tokens, 1M context, the first Haiku with an adjustable effort parameter — and a 100K prompt cutoff that changes the economics.
Haiku 5.5 at a glance
Claude Haiku 5.5 (claude-haiku-5-5) launched October 7, 2026 at $0.10 per million input tokens and $0.50 per million output — the same base rate as GPT-6 Luna, and Anthropic says roughly 75% cheaper to run than Haiku 4.5. It is the first Haiku with an adjustable effort parameter (adaptive thinking, default medium), and it beats Haiku 4.5 on every benchmark Anthropic published: 1620 vs 735 on GDPval-AA v2.1, 39.2% vs 0.0% on Terminal-Bench 4.0. The catch is prompt size — above 100,000 prompt tokens the rate jumps 5x to $0.50/$2.50, which stops undercutting GPT-6 Luna on very long prompts. 1M context, 128K max output, June 2026 cutoff.
Claude Haiku 5.5 specs & price table
Base rates and specs from Anthropic’s pricing page and model docs, verified October 8, 2026. Prompts above 100,000 tokens bill the higher tier shown in the next section.
| Spec | Claude Haiku 5.5 |
|---|---|
| Released | 2026-10-07 |
| API model ID | claude-haiku-5-5 |
| Input (per 1M tokens, ≤100K prompt) | $0.10 |
| Output (per 1M tokens, ≤100K prompt) | $0.50 |
| Cache hit | $0.01 |
| Cache write (5-minute) | $0.125 |
| Cache write (1-hour) | $0.20 |
| Batch (in / out) | $0.05 / $0.25 |
| Context window | 1M |
| Max output (Batch beta up to 300K) | 128K |
| Knowledge cutoff | June 2026 |
| Thinking | Adaptive thinking (effort-controlled) |
| Default effort | effort parameter, default medium |
| Availability | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
| Tokenizer | Newer tokenizer (Claude 4.7+); the same text is ~30% more tokens than on Haiku 4.5 |
| Data retention | Not stated on the pricing page |
All rates above are read from our MODEL_REGISTRY and verified against Anthropic’s pricing page on 2026-10-08 (USD per million tokens). The Batch API bills 50% of input and output, and Anthropic’s Batch beta accepts up to 300,000 output tokens per request.
The >100K prompt cliff: 5x the base rate
Haiku 5.5 is the first Anthropic model on this site with published tiered pricing. The base tier applies up to 100,000 prompt tokens; above that every rate jumps 5x — cache reads included.
| Rate | Up to 100K prompt tokens | Above 100K prompt tokens |
|---|---|---|
| Multiplier vs base | 1x | 5x |
| Input (per 1M tokens) | $0.10 | $0.50 |
| Output (per 1M tokens) | $0.50 | $2.50 |
| Cache hit | $0.01 | $0.05 |
| Cache write (5-minute) | $0.125 | $0.625 |
| Cache write (1-hour) | $0.20 | $1.00 |
Up to 100K prompt tokens Haiku 5.5 bills the base rate; above that the input rate jumps 5x to $0.50 per million tokens, while GPT-6 Luna holds $0.10 to 272K prompt tokens. So between 100K and 272K Haiku 5.5 costs about 5x Luna on input; past 272K Luna doubles to $0.20 while Haiku stays on its 5x tier, and Luna becomes the cheaper of the two on very long prompts. About 90% of Haiku 4.5 requests fall under the cutoff, so most workloads never reach the higher tier.
Haiku 5.5 benchmarks vs Haiku 4.5, GPT-6 Luna and Sonnet 5.5
Seven evaluations from Anthropic's October 7, 2026 launch post. Vendor-reported and compiled from several suites — the Haiku 5.5 System Card documents each row's methodology. Treat them as Anthropic's numbers, not independent results.
| Benchmark | Claude Haiku 5.5 | Claude Haiku 4.5 | GPT-6 Luna | Claude Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (knowledge work) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1, offline subset (computer use) | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity’s Last Exam (no tools / with tools) | 45.9% / 57.4% | 10.2% / 18.7% | Not published | 56.9% / 64.5% |
| Terminal-Bench 4.0 (agentic coding) | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 Main | 46.4% | Not published | 42.4% | 52.1% (Xhigh) |
| Chartography (visual reasoning, no tools) | 46.4% | 6.4% | 29.1% | 61.6% |
Source: Anthropic’s Claude Haiku 5.5 launch post, 2026-10-07 (vendor-reported, not independently verified). Haiku 5.5 improves on Haiku 4.5 in every published row — most sharply on Terminal-Bench 4.0 (39.2% vs 0.0%) and OSWorld 2.1 (72.4% vs 15.7%). Sonnet 5.5 still leads every row where both are published; GPT-6 Luna sits between them.
Claude Haiku 5.5 vs Haiku 4.5: what the upgrade changes
Same job, much faster — but the sticker discount is not the whole bill.
Base rate cut from $1/$5 to $0.10/$0.50
Haiku 4.5 listed $1 in / $5 out per million tokens; Haiku 5.5 lists $0.10/$0.50 — a 10x cut on input. Cache reads fall from $0.10 to $0.01, and the 5-minute cache write from $1.25 to $0.125. Batch is half price at $0.05/$0.25.
Same text, ~30% more tokens
The newer tokenizer shared with Claude 4.7 and later makes identical text count as about 30% more tokens than on Haiku 4.5. Budget for that before you switch a high-volume pipeline on sticker price alone.
Not a small step on benchmarks
Vendor-reported gains run from 1620 vs 735 on GDPval-AA v2.1 to 39.2% vs 0.0% on Terminal-Bench 4.0. Charts are the other outlier: 46.4% vs 6.4% on Chartography.
200K to 1M context
Haiku 4.5 carried a 200K window; Haiku 5.5 raises it to 1M with the same 128K max output, so long documents and agent transcripts fit without chunking — though prompts over 100K tokens bill the 5x tier.
Anthropic puts Haiku 5.5 at roughly 75% cheaper to run than Haiku 4.5, and the benchmark gaps back it up. The catch is the tokenizer: Haiku 5.5 uses the same newer tokenizer as Claude 4.7 and later, so the same text counts as about 30% more tokens than it did on Haiku 4.5. A ~75% cheaper sticker rate is therefore not a 75% cheaper bill. As of 2026-10-08, the base rates are $0.10/$0.50 per million tokens, well below Haiku 4.5 for most workloads.
Adjustable effort: a first for Haiku
Adaptive thinking with the effort parameter, defaulting to medium — the control Haiku 4.5 never had.
The first Haiku with adjustable effort
Earlier Haiku models exposed no effort control. Haiku 5.5 does, which matters most when one model handles both cheap classification and agentic work.
Default effort: medium
Requests run at medium unless you set the parameter explicitly. Anthropic’s docs list effort as the supported way to trade quality against latency and cost.
Adaptive thinking
Thinking is adaptive rather than always-on: the model decides how much to reason per request, and effort bounds that budget.
When to move the dial
Use low effort for high-volume, latency-sensitive jobs (classification, routing, extraction, context compaction) and raise it for multi-step coding or research. Re-benchmark at your chosen effort rather than assuming the launch-post numbers carry over.
Anthropic calls Haiku 5.5 the first Haiku with an adjustable effort parameter. Combined with adaptive thinking, that turns a fixed-cost classifier into a dial: low effort for routing, extraction and compaction; higher effort when a task needs real reasoning. Default is medium.
Where Haiku 5.5 is available
Model IDs and platforms, verified October 8, 2026.
Claude API
Model ID claude-haiku-5-5. Base rates $0.10 in / $0.50 out per million tokens, cache hits $0.01, cache writes $0.125 (5-minute) / $0.20 (1-hour), Batch $0.05/$0.25.
Amazon Bedrock
Available as anthropic.claude-haiku-5-5, with the same tiered rates.
Google Cloud, Microsoft Foundry, Claude Platform on AWS
Anthropic lists Haiku 5.5 on all three alongside first-party and Bedrock access; check each platform’s model list for the exact identifier.
Batch API
Batch bills 50% of input and output and, in Anthropic’s beta, accepts up to 300,000 output tokens per request.
This page was last reviewed 2026-10-08. Model rates, tier thresholds and platform availability move fast — Anthropic’s pricing page and the Haiku 5.5 model docs remain the source of record.
Estimate your Haiku 5.5 bill
Haiku 5.5 is cheap until a prompt crosses 100K tokens, where the rate jumps 5x. Run your real prompt sizes and cache-hit rate through the cost calculator before you commit.