Haiku 5.5 shipped October 7, 2026
This page was last verified on October 8, 2026 against Anthropic's Haiku 5.5 launch post and platform pricing docs, and OpenAI's docs for GPT-6 Luna. Prices and tiers move — re-check the vendor pages before you commit to a bill.
Claude Haiku 5.5 specs and pricingVerified as of October 8, 2026
Claude Haiku 5.5 vs GPT-6 Luna
Two cheap small models on the same sticker price — and a decisive split above 100K prompt tokens. Haiku 5.5 leads Anthropic's own head-to-head benchmarks; Luna undercuts it on long prompts and lives entirely inside OpenAI's stack. Here is the verdict, with every figure dated and sourced.
TL;DR — which one to pick
For most short-prompt work the two cost the same, so pick by ecosystem: Haiku 5.5 if you are in the Claude API/Bedrock stack or want the stronger scores, GPT-6 Luna if you are already on OpenAI's Responses API. Above 100K prompt tokens the price story flips: Haiku 5.5 bills $0.50 in / $2.50 out, while Luna keeps $0.10/$0.50 until 272K — a 200K-token request costs $0.1250 on Haiku 5.5 and $0.0250 on Luna. Above 272K both step up, and Luna stays cheaper. The benchmarks below are Anthropic's own, so treat them as vendor-reported.
Specs and price, side by side
Both models list the same base rate. The difference is what happens to that rate as the prompt grows.
| Item | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Model ID | claude-haiku-5-5 | gpt-6-luna |
| Released | October 7, 2026 | September 22, 2026 |
| Input, base tier (per 1M tokens) | $0.10 | $0.10 |
| Output, base tier (per 1M tokens) | $0.50 | $0.50 |
| Above the base tier | Above 100K prompt tokens: $0.50 in / $2.50 out — 5x the base rate. Nothing extra below the threshold. | Flat $0.10 in / $0.50 out up to 272K prompt tokens; above that, 2x input and cache and 1.5x output. |
| Cache read, base tier | $0.01 | $0.01 |
| Cache write, 5 minutes, base tier | $0.125 | $0.125 |
| Cache write, 1 hour | $0.20 | Not published |
| Batch API (per 1M tokens) | $0.05 / $0.25 | $0.05 / $0.25 |
| Context window | 1M | 1.05M |
| Max output | 128K | 128K |
| Knowledge cutoff | June 2026 | May 18, 2026 |
| Thinking / effort control | Adaptive thinking, default effort medium — the first Haiku with an effort dial | Reasoning model; reasoning effort selectable, fast mode at 2x |
| Where you can call it | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS | OpenAI Responses API |
Haiku 5.5 rates are Anthropic list prices from the October 7, 2026 launch post and platform docs: up to 100K prompt tokens at the base tier, 5x above it. GPT-6 Luna rates are OpenAI's published list prices, verified 2026-09-25: $0.10/$0.50 flat, with prompts above 272K billed at 2x input and cache plus 1.5x output. Cache reads, 5-minute cache writes and Batch rates are identical on the two base tiers; only Haiku 5.5 publishes a 1-hour cache-write rate. All figures are vendor list prices in USD per million tokens.
The decisive part: long prompts flip the cost story
Both models list $0.10 in / $0.50 out per million tokens. Haiku 5.5's base tier ends at 100K prompt tokens; Luna's flat rate survives to 272K. Run the same request through both and the cheaper model changes at 100K.
| Item | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| 90K prompt + 10K output | $0.0140 | $0.0140 |
| 200K prompt + 10K output | $0.1250 | $0.0250 |
| 300K prompt + 10K output | $0.1750 | $0.0675 |
| 300K cached prompt + 10K output | $0.0400 | $0.0135 |
Under 100K prompt tokens: a tie
Both bill $0.10 in / $0.50 out, so a 90K prompt with 10K output costs $0.0140 either way. Choose on ecosystem and benchmarks, not price.
100K–272K prompt tokens: Luna wins by 5x
Haiku 5.5 crosses into its high tier at 100,000 prompt tokens while Luna stays flat, so a 200K prompt with 10K output bills $0.1250 on Haiku 5.5 and $0.0250 on Luna. This is the widest gap between the two models.
Above 272K prompt tokens: Luna still wins
Luna's 2x input / 1.5x output surcharge starts above 272K, but Haiku 5.5 is already at 5x base. For 300K in + 10K out: $0.1750 on Haiku 5.5 against $0.0675 on Luna, roughly 2.6x. Cache reads keep the same multiple in this band.
Costs are list price for a fresh (uncached) prompt: prompt tokens x input rate + output tokens x output rate, in USD, rounded to four decimals. Haiku 5.5 uses its up-to-100K tier ($0.10/$0.50) below 100,000 prompt tokens and its 100K-plus tier ($0.50/$2.50) above. Luna uses $0.10/$0.50 up to 272,000 prompt tokens and 2x input / 1.5x output above. The cached row applies each model's cache-read rate to the whole prompt. No Batch discount, no fast mode.
Head-to-head benchmarks (vendor-reported)
Anthropic ran these seven comparisons in its Haiku 5.5 launch post and published all of them except a Luna figure for Humanity's Last Exam. Higher is better.
| Item | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| GDPval-AA v2.1 | 1620 | 1437 |
| AA-Briefcase v1.1 | 1578 | 1336 |
| OSWorld 2.1 (offline subset) | 72.4% | 48.9% |
| Terminal-Bench 4.0 | 39.2% | 16.4% |
| FrontierCode 1.1 Main | 46.4% | 42.4% |
| Chartography | 46.4% | 29.1% |
| Humanity's Last Exam | 45.9% no tools / 57.4% with tools | Not published — Anthropic's table shows no Luna figure |
Scores come from Anthropic's October 7, 2026 Haiku 5.5 launch post, compiling several evaluation suites; each row's methodology is in the Haiku 5.5 System Card, not the launch page. Vendor-compiled, not an independent head-to-head — no third party has re-run the pair — so read cross-row comparisons with care and check the card before quoting. Haiku 5.5's HLE scores are published (45.9% without tools, 57.4% with tools); GPT-6 Luna's is not.
Ecosystem and platform reach
The rates are close enough that where a model runs often decides the bill.
Haiku 5.5 runs anywhere Claude does
Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS — one model ID to move between them. If your stack is already Claude, this is the cheaper integration.
Luna is OpenAI Responses API only
GPT-6 Luna is served through OpenAI's Responses API. That is one mature, well-trodden surface, but it does not give you Bedrock, Vertex or Foundry as alternates.
Batch and cache rules are near-identical
Both charge $0.05 in / $0.25 out on the Batch API and $0.01 for a base-tier cache read. Switching between them is mostly an SDK change, not a re-architecture.
Which one to pick
Route by prompt length first, then by stack.
Short prompts (under 100K): decide on ecosystem and quality
Same price. Haiku 5.5 holds the stronger vendor-reported scores, so take it if you are in the Claude stack or you care about the benchmark spread.
Long prompts (100K–272K): take Luna
This is the band where the two models genuinely diverge on price — 5x at 200K prompt tokens. Unless Haiku 5.5's accuracy is worth 5x to you, Luna wins here.
Very long prompts (above 272K): still Luna, smaller margin
Luna's surcharge raises its rate above 272K, but it stays roughly 2.5–3.3x cheaper than Haiku 5.5 in that band.
Agentic and computer-use work: Haiku 5.5
Anthropic's own Terminal-Bench 4.0, OSWorld 2.1 and AA-Briefcase comparisons put Haiku 5.5 well ahead, and it is designed to run as a subagent under Opus 5.5 or Sonnet 5.5.
High-volume routing and extraction: either, on price
Under 100K prompt tokens the rates match and so does the Batch discount. Pick the vendor you already pay; Haiku 5.5 also gives you the effort dial to make thinking cheaper.
Re-check before you commit
Both vendors changed small-model pricing twice in 2026. Verify against the linked docs on the day the workload goes live, not when this page was written.
This is a list-price comparison. Volume discounts, committed-spend rates, regional surcharges (Luna charges +10% on regional endpoints) and fast mode (Luna, 2x) are not modelled here.
Compare the bill for your own workload
Put your prompt size, output size and cache-hit rate into the calculator, then read the Haiku 5.5 page for the full spec sheet.