Flash 三者比較 · 2026年9月
Qwen 3.8 Flash vs GLM 5.3 Flash vs DeepSeek V4 Flash
2026年9月時点で最速の中国 lab commodity モデル3つ — 価格、スループット、coding 品質、cheap flash routing が frontier に勝つ条件。
TL;DR
GLM 5.3 Flash が最安($0.10/$0.20)MIT open weights。Qwen 3.8 Flash が flash tier で SWE-bench 58.2% 先行。DeepSeek V4 Flash が raw latency 140ms p50。hard agent work は Fable 5.1 へ。
スペック概要
100万トークンあたり USD、2026年9月 vendor リスト。
| 指標 | Qwen 3.8 Flash | GLM 5.3 Flash | DeepSeek V4 Flash |
|---|---|---|---|
| モデル ID | Qwen-3.8-Flash | GLM-5.3-Flash | DeepSeek-V4-Flash |
| 入力価格(100万トークン) | $0.12 | $0.10 | $0.14 |
| 出力価格(100万トークン) | $0.24 | $0.20 | $0.28 |
| Prompt cache read(100万トークン) | $0.02 | $0.02 | $0.02 |
| コンテキスト | 1M | 1M | 1M |
| 最大出力 | 64K | 64K | 64K |
| ライセンス | Commercial API(Alibaba Cloud) | Open weights — MIT | Open weights — MIT |
Qwen 3.8 Flash は Alibaba Cloud(commercial API)。GLM 5.3 Flash と DeepSeek V4 Flash は MIT open weights、self-host 可。
ベンチマーク比較
Flash tier は frontier reasoning を speed/cost と trade — production 前に honest 比較。
| 指標 | Qwen 3.8 Flash | GLM 5.3 Flash | DeepSeek V4 Flash |
|---|---|---|---|
| SWE-bench Verified | 58.2% | 52.4% | 49.8% |
| HumanEval | 91.4% | 88.6% | 87.2% |
| Latency p50(TTFT) | 180ms | 165ms | 140ms |
| スループット | 420 tok/s | 480 tok/s | 510 tok/s |
SWE-bench/HumanEval は vendor 自行報告、2026年9月。latency は regional edge 測定 — 環境で変動。
flash_models_page.premium_title
flash_models_page.premium_subtitle
Claude Fable 5.1 に escalate するタイミング
Flash は bulk routing、draft、簡単 coding。SWE-bench gap が効くとき — Fable 5.1 $10.00/$50.00/M で 87.7% SWE-bench(87.7%)、flash tier は約50%。hard 10–20% を upstream。
選び方
デフォルト flash lane を選び、失敗で escalate — 全部 frontier 価格で回さない。
高 volume bulk routing
DeepSeek V4 Flash 510 tok/s、$0.14/$0.28 — 分類、要約、簡単 codegen の scale。
中国 hosted commercial API
Qwen 3.8 Flash on Alibaba Cloud、RMB billing、mainland data residency。SLA 必要で self-host しない場合。
open weights self-host
GLM 5.3 Flash(MIT)と DeepSeek V4 Flash(MIT)は modest GPU fleet で可。GLM 5.3 Flash が最安 $0.10/$0.20。
最低リスト価格優先
GLM 5.3 Flash $0.10/$0.20 が Qwen/DeepSeek 入力より安い。無差別 batch はここから。
hard agentic coding は escalate
multi-file 自律 agent、complex refactor、frontier debug は Fable 5.1 — SWE-bench 50–58% cap の flash tier ではない。
月1億 token なら $0.10 と $10 入力差は real — agent 失敗リトライも real。