Released October 7, 2026 — vendor figures verified October 8, 2026

Claude Haiku 5.5 Benchmarks & Pricing

Anthropic’s cheapest and fastest tier, rebuilt: $0.10/$0.50 per million tokens, 1M context, the first Haiku with an adjustable effort parameter — and a 100K prompt cutoff that changes the economics.

Haiku 5.5 at a glance

Claude Haiku 5.5 (claude-haiku-5-5) launched October 7, 2026 at $0.10 per million input tokens and $0.50 per million output — the same base rate as GPT-6 Luna, and Anthropic says roughly 75% cheaper to run than Haiku 4.5. It is the first Haiku with an adjustable effort parameter (adaptive thinking, default medium), and it beats Haiku 4.5 on every benchmark Anthropic published: 1620 vs 735 on GDPval-AA v2.1, 39.2% vs 0.0% on Terminal-Bench 4.0. The catch is prompt size — above 100,000 prompt tokens the rate jumps 5x to $0.50/$2.50, which stops undercutting GPT-6 Luna on very long prompts. 1M context, 128K max output, June 2026 cutoff.

Claude Haiku 5.5 specs & price table

Base rates and specs from Anthropic’s pricing page and model docs, verified October 8, 2026. Prompts above 100,000 tokens bill the higher tier shown in the next section.

SpecClaude Haiku 5.5
Released2026-10-07
API model IDclaude-haiku-5-5
Input (per 1M tokens, ≤100K prompt)$0.10
Output (per 1M tokens, ≤100K prompt)$0.50
Cache hit$0.01
Cache write (5-minute)$0.125
Cache write (1-hour)$0.20
Batch (in / out)$0.05 / $0.25
Context window1M
Max output (Batch beta up to 300K)128K
Knowledge cutoffJune 2026
ThinkingAdaptive thinking (effort-controlled)
Default efforteffort parameter, default medium
AvailabilityClaude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS
TokenizerNewer tokenizer (Claude 4.7+); the same text is ~30% more tokens than on Haiku 4.5
Data retentionNot stated on the pricing page

All rates above are read from our MODEL_REGISTRY and verified against Anthropic’s pricing page on 2026-10-08 (USD per million tokens). The Batch API bills 50% of input and output, and Anthropic’s Batch beta accepts up to 300,000 output tokens per request.

The >100K prompt cliff: 5x the base rate

Haiku 5.5 is the first Anthropic model on this site with published tiered pricing. The base tier applies up to 100,000 prompt tokens; above that every rate jumps 5x — cache reads included.

RateUp to 100K prompt tokensAbove 100K prompt tokens
Multiplier vs base1x5x
Input (per 1M tokens)$0.10$0.50
Output (per 1M tokens)$0.50$2.50
Cache hit$0.01$0.05
Cache write (5-minute)$0.125$0.625
Cache write (1-hour)$0.20$1.00

Up to 100K prompt tokens Haiku 5.5 bills the base rate; above that the input rate jumps 5x to $0.50 per million tokens, while GPT-6 Luna holds $0.10 to 272K prompt tokens. So between 100K and 272K Haiku 5.5 costs about 5x Luna on input; past 272K Luna doubles to $0.20 while Haiku stays on its 5x tier, and Luna becomes the cheaper of the two on very long prompts. About 90% of Haiku 4.5 requests fall under the cutoff, so most workloads never reach the higher tier.

Haiku 5.5 benchmarks vs Haiku 4.5, GPT-6 Luna and Sonnet 5.5

Seven evaluations from Anthropic's October 7, 2026 launch post. Vendor-reported and compiled from several suites — the Haiku 5.5 System Card documents each row's methodology. Treat them as Anthropic's numbers, not independent results.

BenchmarkClaude Haiku 5.5Claude Haiku 4.5GPT-6 LunaClaude Sonnet 5.5
GDPval-AA v2.1 (knowledge work)162073514371840
AA-Briefcase v1.1157861413361824
OSWorld 2.1, offline subset (computer use)72.4%15.7%48.9%83.9%
Humanity’s Last Exam (no tools / with tools)45.9% / 57.4%10.2% / 18.7%Not published56.9% / 64.5%
Terminal-Bench 4.0 (agentic coding)39.2%0.0%16.4%70.6%
FrontierCode 1.1 Main46.4%Not published42.4%52.1% (Xhigh)
Chartography (visual reasoning, no tools)46.4%6.4%29.1%61.6%

Source: Anthropic’s Claude Haiku 5.5 launch post, 2026-10-07 (vendor-reported, not independently verified). Haiku 5.5 improves on Haiku 4.5 in every published row — most sharply on Terminal-Bench 4.0 (39.2% vs 0.0%) and OSWorld 2.1 (72.4% vs 15.7%). Sonnet 5.5 still leads every row where both are published; GPT-6 Luna sits between them.

Claude Haiku 5.5 vs Haiku 4.5: what the upgrade changes

Same job, much faster — but the sticker discount is not the whole bill.

  • Base rate cut from $1/$5 to $0.10/$0.50

    Haiku 4.5 listed $1 in / $5 out per million tokens; Haiku 5.5 lists $0.10/$0.50 — a 10x cut on input. Cache reads fall from $0.10 to $0.01, and the 5-minute cache write from $1.25 to $0.125. Batch is half price at $0.05/$0.25.

  • Same text, ~30% more tokens

    The newer tokenizer shared with Claude 4.7 and later makes identical text count as about 30% more tokens than on Haiku 4.5. Budget for that before you switch a high-volume pipeline on sticker price alone.

  • Not a small step on benchmarks

    Vendor-reported gains run from 1620 vs 735 on GDPval-AA v2.1 to 39.2% vs 0.0% on Terminal-Bench 4.0. Charts are the other outlier: 46.4% vs 6.4% on Chartography.

  • 200K to 1M context

    Haiku 4.5 carried a 200K window; Haiku 5.5 raises it to 1M with the same 128K max output, so long documents and agent transcripts fit without chunking — though prompts over 100K tokens bill the 5x tier.

Anthropic puts Haiku 5.5 at roughly 75% cheaper to run than Haiku 4.5, and the benchmark gaps back it up. The catch is the tokenizer: Haiku 5.5 uses the same newer tokenizer as Claude 4.7 and later, so the same text counts as about 30% more tokens than it did on Haiku 4.5. A ~75% cheaper sticker rate is therefore not a 75% cheaper bill. As of 2026-10-08, the base rates are $0.10/$0.50 per million tokens, well below Haiku 4.5 for most workloads.

Adjustable effort: a first for Haiku

Adaptive thinking with the effort parameter, defaulting to medium — the control Haiku 4.5 never had.

  • The first Haiku with adjustable effort

    Earlier Haiku models exposed no effort control. Haiku 5.5 does, which matters most when one model handles both cheap classification and agentic work.

  • Default effort: medium

    Requests run at medium unless you set the parameter explicitly. Anthropic’s docs list effort as the supported way to trade quality against latency and cost.

  • Adaptive thinking

    Thinking is adaptive rather than always-on: the model decides how much to reason per request, and effort bounds that budget.

  • When to move the dial

    Use low effort for high-volume, latency-sensitive jobs (classification, routing, extraction, context compaction) and raise it for multi-step coding or research. Re-benchmark at your chosen effort rather than assuming the launch-post numbers carry over.

Anthropic calls Haiku 5.5 the first Haiku with an adjustable effort parameter. Combined with adaptive thinking, that turns a fixed-cost classifier into a dial: low effort for routing, extraction and compaction; higher effort when a task needs real reasoning. Default is medium.

Where Haiku 5.5 is available

Model IDs and platforms, verified October 8, 2026.

  • Claude API

    Model ID claude-haiku-5-5. Base rates $0.10 in / $0.50 out per million tokens, cache hits $0.01, cache writes $0.125 (5-minute) / $0.20 (1-hour), Batch $0.05/$0.25.

  • Amazon Bedrock

    Available as anthropic.claude-haiku-5-5, with the same tiered rates.

  • Google Cloud, Microsoft Foundry, Claude Platform on AWS

    Anthropic lists Haiku 5.5 on all three alongside first-party and Bedrock access; check each platform’s model list for the exact identifier.

  • Batch API

    Batch bills 50% of input and output and, in Anthropic’s beta, accepts up to 300,000 output tokens per request.

This page was last reviewed 2026-10-08. Model rates, tier thresholds and platform availability move fast — Anthropic’s pricing page and the Haiku 5.5 model docs remain the source of record.

Estimate your Haiku 5.5 bill

Haiku 5.5 is cheap until a prompt crosses 100K tokens, where the rate jumps 5x. Run your real prompt sizes and cache-hit rate through the cost calculator before you commit.

Frequently asked questions