Sourced from Anthropic's June 9, 2026 launch table
Claude Fable 5 benchmarks
The full published table — 80.3% SWE-Bench Pro, the top public-model score — plus the honest read the launch posts bury: which starred rows belong to Mythos 5, and why safeguard-flagged domains score at Opus 4.8 level.
Benchmarks — with the honest footnote
Anthropic's published results put Claude Fable5 ahead on agentic coding and knowledge work. Here's the full table, and the caveat the launch posts bury.
| Benchmark | Claude Fable5 / Mythos 5 | Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|
| SWE-Bench Pro (coding) | 80.3% | 69.2% | 58.6% | 54.2% |
| FrontierCode (Diamond) | 29.3% | 13.4% | 5.7% | — |
| GDPval-AA (knowledge work) | 1932 | 1890 | 1769 | 1314 |
| GDP.pdf (vision, no tools) | 29.8% | 22.5% | 24.9% | 16.7% |
| OSWorld-Verified (computer use) | 85.0% | 83.4% | 78.7% | 76.2% |
On cybersecurity and biology evaluations (e.g. ExploitBench), Claude Fable5's safeguards cap its scores near Opus 4.8 — the published top scores there belong to the unrestricted Claude Mythos 5. Source: Anthropic announcement, June 9, 2026.
How to read these numbers honestly
- The starred rows in Anthropic's table — cybersecurity and biology evaluations like ExploitBench — are Claude Mythos 5 scores, not Fable 5. They are real, but you almost certainly cannot reproduce them.
- In safeguard-flagged domains, Fable 5's own results land near Opus 4.8 — because when a classifier fires, Opus 4.8 is literally the model answering. Since the July 1 restoration, some routine coding requests also temporarily fall back to Opus while the new classifiers are tuned.
- Coding, knowledge work, and computer-use scores are genuine Fable 5 results — for those, what is published is what you get.
Rule of thumb: if a domain is restricted enough to trigger Fable 5's safeguards, the published top score belongs to Mythos 5 — and the score you actually experience belongs to Opus 4.8.
Against the rest of the field
Fable 5 leads GPT-5.5 by 22 points on SWE-Bench Pro and 5x on FrontierCode, and Gemini 3.1 Pro by 26 points on SWE-Bench Pro — at roughly 4–8x the price. Sonnet 5 ($2/$10 intro) covers most day-to-day coding at close to Opus quality. Full breakdowns:
Benchmark FAQ
Go deeper
Build on Claude Fable5 with your eyes open
Know the Claude Fable5 cost math, the safeguard behavior, and the right tier before you commit. Create a free account to get notified when specs, prices, or access channels change.