Prices verified as of September 2026
All three flash models priced from vendor list rates as of September 2026. Benchmarks vendor-reported unless noted.
Open the cost calculatorFlash models three-way · September 2026
Qwen 3.8 Flash vs GLM 5.3 Flash vs DeepSeek V4 Flash
The three fastest Chinese-lab commodity models as of September 2026 — price, throughput, coding quality, and when cheap flash routing beats frontier models.
TL;DR
GLM 5.3 Flash is cheapest ($0.10/$0.20) with MIT open weights. Qwen 3.8 Flash leads SWE-bench among flash tiers (58.2%). DeepSeek V4 Flash wins raw latency (140ms p50). Escalate hard agent work to Fable 5.1.
Specs at a glance
USD per million tokens, vendor list prices as of September 2026.
| Metric | Qwen 3.8 Flash | GLM 5.3 Flash | DeepSeek V4 Flash |
|---|---|---|---|
| Model ID | Qwen-3.8-Flash | GLM-5.3-Flash | DeepSeek-V4-Flash |
| Input price (per M tokens) | $0.12 | $0.10 | $0.14 |
| Output price (per M tokens) | $0.24 | $0.20 | $0.28 |
| Prompt cache read (per M tokens) | $0.02 | $0.02 | $0.02 |
| Context window | 1M | 1M | 1M |
| Max output | 64K | 64K | 64K |
| License | Commercial API (Alibaba Cloud) | Open weights — MIT | Open weights — MIT |
Qwen 3.8 Flash runs on Alibaba Cloud (commercial API). GLM 5.3 Flash and DeepSeek V4 Flash are MIT open weights and self-hostable.
Benchmark comparison
Flash-tier models trade frontier reasoning for speed and cost — compare honestly before routing production traffic.
| Metric | Qwen 3.8 Flash | GLM 5.3 Flash | DeepSeek V4 Flash |
|---|---|---|---|
| SWE-bench Verified | 58.2% | 52.4% | 49.8% |
| HumanEval | 91.4% | 88.6% | 87.2% |
| Latency p50 (TTFT) | 180ms | 165ms | 140ms |
| Throughput | 420 tok/s | 480 tok/s | 510 tok/s |
SWE-bench and HumanEval figures vendor-reported as of September 2026. Latency measured at regional edge nodes — your mileage varies.
flash_models_page.premium_title
flash_models_page.premium_subtitle
When to escalate to Claude Fable 5.1
Flash models handle bulk routing, drafts, and simple coding. When SWE-bench gaps matter — Fable 5.1 at $10.00/$50.00 per M tokens hits 87.7% SWE-bench (87.7%) vs ~50% on flash tiers. Route the hard 10–20% upstream.
How to decide
Pick a default flash lane, then escalate failures — don't run everything on frontier pricing.
High-volume bulk routing
DeepSeek V4 Flash at 510 tok/s and $0.14/$0.28 suits classification, summarization, and simple codegen at scale.
China-hosted commercial API
Qwen 3.8 Flash on Alibaba Cloud with RMB billing and mainland data residency. Best when you need vendor SLA without self-hosting.
Self-hosting with open weights
GLM 5.3 Flash (MIT) and DeepSeek V4 Flash (MIT) run on modest GPU fleets. GLM 5.3 Flash is the cheapest list price at $0.10/$0.20.
Lowest list price wins
GLM 5.3 Flash at $0.10/$0.20 undercuts Qwen and DeepSeek on input. For undifferentiated batch work, start here.
Hard agentic coding needs escalation
Multi-file autonomous agents, complex refactors, and frontier debugging belong on Fable 5.1 — not flash tiers capped around 50–58% SWE-bench.
At 100M tokens/month, the gap between $0.10 and $10 input is real — but so is the cost of failed agent retries.
Model your real costs
Compare flash routing vs Fable 5.1 escalation with your actual token mix.