Haiku 5.5 shipped October 7, 2026

This page was last verified on October 8, 2026 against Anthropic's Haiku 5.5 launch post and platform pricing docs, and OpenAI's docs for GPT-6 Luna. Prices and tiers move — re-check the vendor pages before you commit to a bill.

Claude Haiku 5.5 specs and pricing

Verified as of October 8, 2026

Claude Haiku 5.5 vs GPT-6 Luna

Two cheap small models on the same sticker price — and a decisive split above 100K prompt tokens. Haiku 5.5 leads Anthropic's own head-to-head benchmarks; Luna undercuts it on long prompts and lives entirely inside OpenAI's stack. Here is the verdict, with every figure dated and sourced.

TL;DR — which one to pick

For most short-prompt work the two cost the same, so pick by ecosystem: Haiku 5.5 if you are in the Claude API/Bedrock stack or want the stronger scores, GPT-6 Luna if you are already on OpenAI's Responses API. Above 100K prompt tokens the price story flips: Haiku 5.5 bills $0.50 in / $2.50 out, while Luna keeps $0.10/$0.50 until 272K — a 200K-token request costs $0.1250 on Haiku 5.5 and $0.0250 on Luna. Above 272K both step up, and Luna stays cheaper. The benchmarks below are Anthropic's own, so treat them as vendor-reported.

Specs and price, side by side

Both models list the same base rate. The difference is what happens to that rate as the prompt grows.

ItemClaude Haiku 5.5GPT-6 Luna
Model IDclaude-haiku-5-5gpt-6-luna
ReleasedOctober 7, 2026September 22, 2026
Input, base tier (per 1M tokens)$0.10$0.10
Output, base tier (per 1M tokens)$0.50$0.50
Above the base tierAbove 100K prompt tokens: $0.50 in / $2.50 out — 5x the base rate. Nothing extra below the threshold.Flat $0.10 in / $0.50 out up to 272K prompt tokens; above that, 2x input and cache and 1.5x output.
Cache read, base tier$0.01$0.01
Cache write, 5 minutes, base tier$0.125$0.125
Cache write, 1 hour$0.20Not published
Batch API (per 1M tokens)$0.05 / $0.25$0.05 / $0.25
Context window1M1.05M
Max output128K128K
Knowledge cutoffJune 2026May 18, 2026
Thinking / effort controlAdaptive thinking, default effort medium — the first Haiku with an effort dialReasoning model; reasoning effort selectable, fast mode at 2x
Where you can call itClaude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWSOpenAI Responses API

Haiku 5.5 rates are Anthropic list prices from the October 7, 2026 launch post and platform docs: up to 100K prompt tokens at the base tier, 5x above it. GPT-6 Luna rates are OpenAI's published list prices, verified 2026-09-25: $0.10/$0.50 flat, with prompts above 272K billed at 2x input and cache plus 1.5x output. Cache reads, 5-minute cache writes and Batch rates are identical on the two base tiers; only Haiku 5.5 publishes a 1-hour cache-write rate. All figures are vendor list prices in USD per million tokens.

The decisive part: long prompts flip the cost story

Both models list $0.10 in / $0.50 out per million tokens. Haiku 5.5's base tier ends at 100K prompt tokens; Luna's flat rate survives to 272K. Run the same request through both and the cheaper model changes at 100K.

ItemClaude Haiku 5.5GPT-6 Luna
90K prompt + 10K output$0.0140$0.0140
200K prompt + 10K output$0.1250$0.0250
300K prompt + 10K output$0.1750$0.0675
300K cached prompt + 10K output$0.0400$0.0135
  • Under 100K prompt tokens: a tie

    Both bill $0.10 in / $0.50 out, so a 90K prompt with 10K output costs $0.0140 either way. Choose on ecosystem and benchmarks, not price.

  • 100K–272K prompt tokens: Luna wins by 5x

    Haiku 5.5 crosses into its high tier at 100,000 prompt tokens while Luna stays flat, so a 200K prompt with 10K output bills $0.1250 on Haiku 5.5 and $0.0250 on Luna. This is the widest gap between the two models.

  • Above 272K prompt tokens: Luna still wins

    Luna's 2x input / 1.5x output surcharge starts above 272K, but Haiku 5.5 is already at 5x base. For 300K in + 10K out: $0.1750 on Haiku 5.5 against $0.0675 on Luna, roughly 2.6x. Cache reads keep the same multiple in this band.

Costs are list price for a fresh (uncached) prompt: prompt tokens x input rate + output tokens x output rate, in USD, rounded to four decimals. Haiku 5.5 uses its up-to-100K tier ($0.10/$0.50) below 100,000 prompt tokens and its 100K-plus tier ($0.50/$2.50) above. Luna uses $0.10/$0.50 up to 272,000 prompt tokens and 2x input / 1.5x output above. The cached row applies each model's cache-read rate to the whole prompt. No Batch discount, no fast mode.

Head-to-head benchmarks (vendor-reported)

Anthropic ran these seven comparisons in its Haiku 5.5 launch post and published all of them except a Luna figure for Humanity's Last Exam. Higher is better.

ItemClaude Haiku 5.5GPT-6 Luna
GDPval-AA v2.116201437
AA-Briefcase v1.115781336
OSWorld 2.1 (offline subset)72.4%48.9%
Terminal-Bench 4.039.2%16.4%
FrontierCode 1.1 Main46.4%42.4%
Chartography46.4%29.1%
Humanity's Last Exam45.9% no tools / 57.4% with toolsNot published — Anthropic's table shows no Luna figure

Scores come from Anthropic's October 7, 2026 Haiku 5.5 launch post, compiling several evaluation suites; each row's methodology is in the Haiku 5.5 System Card, not the launch page. Vendor-compiled, not an independent head-to-head — no third party has re-run the pair — so read cross-row comparisons with care and check the card before quoting. Haiku 5.5's HLE scores are published (45.9% without tools, 57.4% with tools); GPT-6 Luna's is not.

Ecosystem and platform reach

The rates are close enough that where a model runs often decides the bill.

  • Haiku 5.5 runs anywhere Claude does

    Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS — one model ID to move between them. If your stack is already Claude, this is the cheaper integration.

  • Luna is OpenAI Responses API only

    GPT-6 Luna is served through OpenAI's Responses API. That is one mature, well-trodden surface, but it does not give you Bedrock, Vertex or Foundry as alternates.

  • Batch and cache rules are near-identical

    Both charge $0.05 in / $0.25 out on the Batch API and $0.01 for a base-tier cache read. Switching between them is mostly an SDK change, not a re-architecture.

Which one to pick

Route by prompt length first, then by stack.

  • Short prompts (under 100K): decide on ecosystem and quality

    Same price. Haiku 5.5 holds the stronger vendor-reported scores, so take it if you are in the Claude stack or you care about the benchmark spread.

  • Long prompts (100K–272K): take Luna

    This is the band where the two models genuinely diverge on price — 5x at 200K prompt tokens. Unless Haiku 5.5's accuracy is worth 5x to you, Luna wins here.

  • Very long prompts (above 272K): still Luna, smaller margin

    Luna's surcharge raises its rate above 272K, but it stays roughly 2.5–3.3x cheaper than Haiku 5.5 in that band.

  • Agentic and computer-use work: Haiku 5.5

    Anthropic's own Terminal-Bench 4.0, OSWorld 2.1 and AA-Briefcase comparisons put Haiku 5.5 well ahead, and it is designed to run as a subagent under Opus 5.5 or Sonnet 5.5.

  • High-volume routing and extraction: either, on price

    Under 100K prompt tokens the rates match and so does the Batch discount. Pick the vendor you already pay; Haiku 5.5 also gives you the effort dial to make thinking cheaper.

  • Re-check before you commit

    Both vendors changed small-model pricing twice in 2026. Verify against the linked docs on the day the workload goes live, not when this page was written.

This is a list-price comparison. Volume discounts, committed-spend rates, regional surcharges (Luna charges +10% on regional endpoints) and fast mode (Luna, 2x) are not modelled here.

Compare the bill for your own workload

Put your prompt size, output size and cache-hit rate into the calculator, then read the Haiku 5.5 page for the full spec sheet.

FAQ: Haiku 5.5 vs GPT-6 Luna