Verified as of August 4, 2026

Claude Fable 5 — All Model Comparisons

One page to see the whole picture: how Claude Fable 5 compares against Anthropic's own lineup and every serious rival of 2026, on price, agentic coding, context, openness and best use case — with every figure dated and sourced, and links to the full per-model comparisons.

TL;DR — Fable 5 in one paragraph

Claude Fable 5 is Anthropic's highest-published agentic-coding model — 80.3% on SWE-Bench Pro, $10/$50 per million tokens — but it is no longer the default Claude: since July 24, 2026, Claude Opus 5 is the default model on Claude Max, matching Fable 5's capability band at half the price ($5/$25) with no mandatory data retention. Pick Fable 5 for the hardest agentic work — long-horizon autonomy, frontier-hard coding, full-repo 1M-token analysis — and Opus 5 (or a cheaper rival) for everything else. The tables below show all nine alternatives at a glance.

Last verified2026-08-04

Prices and benchmark figures on this page were checked against the cited sources on August 4, 2026. Anthropic figures come from src/lib/model-pricing.ts, which is checked against Anthropic's pricing docs; competitor figures come from the individual comparison pages on this site (each dated and sourced there) plus the vendor pages cited below. Where a figure could not be confirmed it is marked "not confirmed" with the date.

The master comparison matrix

API list prices per million tokens, USD, and the headline facts that decide most choices. Full context and caveats below and on each comparison page.

ModelInput (per M)Output (per M)ContextOpen weightsBest forSource
Claude Fable 5$10.00$50.001MClosed APIFrontier agentic codingOfficial (Anthropic / vendor docs)
Claude Opus 5$5.00$25.001MClosed APIDefault lane, no retentionOfficial (Anthropic / vendor docs)
Claude Sonnet 5$2.00*$10.00*1MClosed APIBalanced Claude tierOfficial (Anthropic / vendor docs)
Gemini 3.1 Pro$2.00$12.001MClosed APIMultimodal, Google stackOfficial (Anthropic / vendor docs)
GPT-5.5$5.00$30.00400KClosed APIOpenAI-stack productsOfficial (Anthropic / vendor docs)
GLM 5.2~$1.40~$4.401MOpen weights (MIT)Free local open weightsVendor-reported
DeepSeek V4 Pro$0.435$0.871MOpen weightsCheap strong reasoningVendor-reported
DeepSeek V4 Flash$0.14$0.281MOpen weights (MIT)High-volume commodityVendor-reported
Qwen 3.8 Max≈$2.00≈$6.001MAnnounced, not releasedChina + priceVendor-reported
Kimi K3$3.00$15.001MOpen weightsChinese-language + open weightsVendor-reported

Sonnet 5 prices are introductory through 2026-08-31. Gemini 3.1 Pro rates apply to prompts up to 200K tokens (higher tiers above); GPT-5.5 rates apply up to 272K, with higher long-context rates above and a cheaper non-reasoning mode at $1.25/$10. GLM 5.2, DeepSeek V4 Pro and Flash, Qwen 3.8 Max and Kimi K3 prices are vendor-reported and not independently verified; Qwen's $2/$6 is an international-list estimate — see the note below the benchmark table. Source labels: "Official" = vendor documentation; "Vendor-reported" = the vendor's own announcement, not independently verified.

Headline benchmark matrix — who measured what

Agentic-coding and knowledge-work scores for the models with published numbers. Every figure is labeled by source; treat any lab-published number as directional.

ModelClaude Fable 5Claude Opus 5GPT-5.5GLM 5.2
SWE-Bench Pro (agentic coding)80.3%Not confirmed (as of 2026-08-04)58.6%62.1%
FrontierCode Diamond (hard coding)29.3%No published score5.7%No published score
GDPval-AA (knowledge work)1932No published score1769No published score
OSWorld-Verified (computer use)85.0%No published score78.7%No published score

Sources: Fable 5, Opus 5, Sonnet 5 figures from Anthropic's June 9, 2026 launch announcement and model docs (vendor self-reported, verified against Anthropic sources August 4, 2026). GPT-5.5 figures from OpenAI's launch (vendor self-reported). GLM 5.2 SWE-Bench Pro 62.1% and the 74.4% FrontierSWE figure are vendor-reported via the GLM 5.2 tech report (June 2026); Opus 4.8 and GPT-5.5 columns in the same table are Anthropic's / OpenAI's figures. Gemini 3.1 Pro, DeepSeek V4 Pro, DeepSeek V4 Flash, Qwen 3.8 Max and Kimi K3 have no published SWE-Bench Pro score as of August 4, 2026 — the cells say so rather than guessing.

Evidence parity: Fable 5's 80.3% and GPT-5.5's 58.6% are the only two SWE-Bench Pro figures that share the same harness definition and scoring protocol, even though the runs used settings chosen by the respective labs. GLM 5.2's 62.1% comes from a different harness implementation and is not directly comparable. GDPval-AA 1932 (Fable 5) and 1769 (GPT-5.5) are Artificial Analysis measurements (independent). No model other than Fable 5 has an independently verified SWE-Bench Pro score as of 2026-08-04. For the full Fable 5 benchmark set, see the benchmarks page.

Decision guide — by budget

What your price ceiling says about which model to start with.

  • You can pay frontier prices ($50+ / M output)

    Fable 5 ($10/$50) for the hardest agentic work; Opus 5 ($5/$25) as the default lane — same capability band on published benchmarks, no 30-day retention. This is the recommended two-tier Claude setup.

  • You want a solid model under ~$15 / M output

    Sonnet 5 ($2/$10, introductory through 2026-08-31), Gemini 3.1 Pro ($2/$12, under 200K-token prompts) or Kimi K3 ($3/$15). Sonnet 5 stays in the Claude ecosystem; Gemini and K3 bring broader multimodal and open-weight options respectively.

  • You're cost-sensitive or high-volume

    GPT-5.5 ($5/$30, or $1.25/$10 in its cheaper non-reasoning mode), GLM 5.2 (~$1.40/$4.40 via Z.ai), DeepSeek V4 Pro ($0.435/$0.87) or DeepSeek V4 Flash ($0.14/$0.28). Flash is the cheapest serious coding model in the market — roughly 70x cheaper than Fable 5 on input.

  • You want open weights or self-hosting

    GLM 5.2 (MIT license, free via Ollama), DeepSeek V4 Pro and V4 Flash (MIT for Flash), or Kimi K3 (~1.4 TB weights, needs ~18 H100-class GPUs — a fleet project, not a laptop model). Qwen 3.8 Max open weights were announced for "next week" on August 3, 2026 but not yet released — not confirmed.

Decision guide — by task type

The honest answer is almost always "both, routed by task." Here is the per-task call.

  • Agentic coding, long-horizon autonomy

    Fable 5 (80.3% SWE-Bench Pro) or Opus 5 — the two highest-published models, with Opus 5 at half the price. GPT-5.5 trails at 58.6%; GLM 5.2 is competitive for mid-complexity single-file work at 62.1%.

  • Everyday and mid-complexity coding

    Any of the mid-tier models clears the bar: Sonnet 5, GPT-5.5, GLM 5.2, DeepSeek V4 Pro or Flash. Pick by price and ecosystem. Fable 5 is overkill here and its premium rarely pays.

  • Long context (1M-token windows)

    Fable 5, Opus 5, Sonnet 5, Gemini 3.1 Pro, GLM 5.2, DeepSeek V4 Pro/Flash, Qwen 3.8 Max and Kimi K3 all offer 1M context. GPT-5.5 tops out at 400K. For context-heavy traffic, the cheaper models win on price; for full-repo reasoning at depth, Fable 5.

  • Chinese-language or China-hosted work

    Kimi K3 (leading Chinese-lab model on BenchAlign v5, China-hosted API) or Qwen 3.8 Max (Alibaba Cloud, RMB billing, ~$2/$6). Fable 5 is a US-export-controlled API served from outside China — a liability if data must stay in China.

  • Images, video or audio input

    Gemini 3.1 Pro (#1 on MMMU-Pro) for the broadest multimodal stack; Qwen 3.8 Max accepts text, image and video. Fable 5 handles image and PDF input, but media-rich workflows are not its edge.

Decision guide — by hard constraint

Some requirements decide before any benchmark does.

  • You must stay on the Anthropic API

    The choice is internal: Fable 5 for the hardest agentic work, Opus 5 as the default lane at half the price, Sonnet 5 for volume and budget. All share the same API surface, caching and tooling.

  • You need zero data retention

    Fable 5's mandatory 30-day retention (zero-data-retention agreements do not apply) rules it out. Opus 5 and Sonnet 5 have no retention requirement; Gemini 3.1 Pro's governance differs and follows Google's terms.

  • You need a model that cannot be switched off

    Fable 5 was suspended worldwide June 12 – July 1, 2026 under a US export-control directive and a repeat is possible in principle. MIT-licensed open weights (GLM 5.2, DeepSeek V4 Flash, soon Qwen 3.8 Max) cannot be revoked — the decisive argument for international and enterprise teams.

  • You need open weights or self-hosting

    GLM 5.2 (MIT), DeepSeek V4 Flash (MIT) or Kimi K3 (Modified MIT, terms not fully confirmed). No Anthropic model is open. Qwen 3.8 Max weights were announced but not yet released as of 2026-08-04.

  • You live in the Google or OpenAI ecosystem

    Gemini 3.1 Pro if you're in Google's stack (Workspace, Vertex, Gemini app free tier); GPT-5.5 if you're in OpenAI's (structured outputs, Assistants). Switching costs often outweigh benchmark gaps.

The decision tree, condensed

Three questions, one answer each — then dive into the full comparison.

Q1 — Budget: over ~$50/M output, you can run Fable 5 or Opus 5 as the escalation lane; under it, start from the mid-tier (Sonnet 5, Gemini 3.1 Pro, GPT-5.5, Kimi K3) or low-tier (GLM 5.2, DeepSeek V4 Pro/Flash, Qwen 3.8 Max) lists. Q2 — Task: frontier-hard agentic coding, long-horizon autonomy, or full-repo 1M analysis point to Fable 5; everything else points to the cheapest model that clears your quality bar. Q3 — Constraints: must stay on Anthropic, must have zero retention, must be open-weight, must be China-hosted — each overrides the first two questions.

The general rule that survives every matrix: Fable 5 earns its premium only on work where a cheap model actually fails and retries (or humans) cost real money. Use the cheapest lane that succeeds, and escalate the hard 10–20% to Fable 5.

Evidence parity — which numbers share a harness

How to read the tables without being misled by inconsistent benchmarks.

  • Same-harness comparisons

    Fable 5 (80.3%) and GPT-5.5 (58.6%) are the only SWE-Bench Pro figures that share the same harness definition and scoring protocol — the only directly comparable row in the table, settings chosen by each lab notwithstanding.

  • Independently measured numbers

    GDPval-AA 1932 (Fable 5) and 1769 (GPT-5.5) are Artificial Analysis measurements. Kimi K3's DeepSWE 68.5% vs Fable 5's 69.9% comes from a Together AI run (August 2026, pass@1) — the only independent head-to-head in this lineup. Everything else is vendor self-reported.

  • DeepSeek's claims

    V4 Flash reports 79.0% on SWE-Bench Verified and 91.6 on LiveCodeBench (official technical report, not independently verified), and its July 31 agentic-upgrade claims have no independent verification. No SWE-Bench Pro figure exists for the Flash or Pro line — not confirmed as of 2026-08-04.

  • Qwen 3.8 Max's claims

    Alibaba claims "second only to Fable 5" on LMArena and the lead on agentic computer use (OSWorld-Verified 86.1, vendor-reported), but shipped no model card or benchmark table at launch — unverified as of 2026-08-04.

  • Kimi K3's claims

    Moonshot self-reports 67.5% on DeepSWE (max reasoning) against Fable 5's 70.0%; an independent Together AI run found them nearly tied at 68.5% vs 69.9% pass@1. K3 has no published SWE-Bench Pro score. License terms not fully confirmed.

Run the numbers on your workload

The cost calculator currently models Anthropic's own lineup (Fable 5, Opus 5, Sonnet 5 and more) with effort levels, caching and the safeguard fallback. It does not yet switch to rival models — for competitor comparisons, use the list prices in the matrix above and the worked examples on each comparison page.

Still deciding?

Read the deepest comparison — Fable 5 vs Opus 5, the two models that decide most Claude decisions — or price out your actual workload.

FAQ