Prices verified as of September 2026

All three flash models priced from vendor list rates as of September 2026. Benchmarks vendor-reported unless noted.

Open the cost calculator

Flash models three-way · September 2026

Qwen 3.8 Flash vs GLM 5.3 Flash vs DeepSeek V4 Flash

The three fastest Chinese-lab commodity models as of September 2026 — price, throughput, coding quality, and when cheap flash routing beats frontier models.

TL;DR

GLM 5.3 Flash is cheapest ($0.10/$0.20) with MIT open weights. Qwen 3.8 Flash leads SWE-bench among flash tiers (58.2%). DeepSeek V4 Flash wins raw latency (140ms p50). Escalate hard agent work to Fable 5.1.

Specs at a glance

USD per million tokens, vendor list prices as of September 2026.

MetricQwen 3.8 FlashGLM 5.3 FlashDeepSeek V4 Flash
Model IDQwen-3.8-FlashGLM-5.3-FlashDeepSeek-V4-Flash
Input price (per M tokens)$0.12$0.10$0.14
Output price (per M tokens)$0.24$0.20$0.28
Prompt cache read (per M tokens)$0.02$0.02$0.02
Context window1M1M1M
Max output64K64K64K
LicenseCommercial API (Alibaba Cloud)Open weights — MITOpen weights — MIT

Qwen 3.8 Flash runs on Alibaba Cloud (commercial API). GLM 5.3 Flash and DeepSeek V4 Flash are MIT open weights and self-hostable.

Benchmark comparison

Flash-tier models trade frontier reasoning for speed and cost — compare honestly before routing production traffic.

MetricQwen 3.8 FlashGLM 5.3 FlashDeepSeek V4 Flash
SWE-bench Verified58.2%52.4%49.8%
HumanEval91.4%88.6%87.2%
Latency p50 (TTFT)180ms165ms140ms
Throughput420 tok/s480 tok/s510 tok/s

SWE-bench and HumanEval figures vendor-reported as of September 2026. Latency measured at regional edge nodes — your mileage varies.

flash_models_page.premium_title

flash_models_page.premium_subtitle

When to escalate to Claude Fable 5.1

Flash models handle bulk routing, drafts, and simple coding. When SWE-bench gaps matter — Fable 5.1 at $10.00/$50.00 per M tokens hits 87.7% SWE-bench (87.7%) vs ~50% on flash tiers. Route the hard 10–20% upstream.

How to decide

Pick a default flash lane, then escalate failures — don't run everything on frontier pricing.

  • High-volume bulk routing

    DeepSeek V4 Flash at 510 tok/s and $0.14/$0.28 suits classification, summarization, and simple codegen at scale.

  • China-hosted commercial API

    Qwen 3.8 Flash on Alibaba Cloud with RMB billing and mainland data residency. Best when you need vendor SLA without self-hosting.

  • Self-hosting with open weights

    GLM 5.3 Flash (MIT) and DeepSeek V4 Flash (MIT) run on modest GPU fleets. GLM 5.3 Flash is the cheapest list price at $0.10/$0.20.

  • Lowest list price wins

    GLM 5.3 Flash at $0.10/$0.20 undercuts Qwen and DeepSeek on input. For undifferentiated batch work, start here.

  • Hard agentic coding needs escalation

    Multi-file autonomous agents, complex refactors, and frontier debugging belong on Fable 5.1 — not flash tiers capped around 50–58% SWE-bench.

At 100M tokens/month, the gap between $0.10 and $10 input is real — but so is the cost of failed agent retries.

Model your real costs

Compare flash routing vs Fable 5.1 escalation with your actual token mix.

FAQ