Цены сверены на сентябрь 2026
Все три flash — vendor list rates на сентябрь 2026. Бенчмарки vendor-reported unless noted.
Калькулятор стоимостиFlash models three-way · сентябрь 2026
Qwen 3.8 Flash vs GLM 5.3 Flash vs DeepSeek V4 Flash
Три fastest Chinese-lab commodity models на сентябрь 2026 — цена, throughput, coding quality, когда cheap flash routing beats frontier.
Кратко
GLM 5.3 Flash cheapest ($0.10/$0.20) с MIT open weights. Qwen 3.8 Flash лидирует SWE-bench среди flash tiers (58.2%). DeepSeek V4 Flash — raw latency 140ms p50. Hard agent work — escalate на Fable 5.1.
Спеки одним взглядом
USD за M tokens, vendor list на сентябрь 2026.
| Метрика | Qwen 3.8 Flash | GLM 5.3 Flash | DeepSeek V4 Flash |
|---|---|---|---|
| Model ID | Qwen-3.8-Flash | GLM-5.3-Flash | DeepSeek-V4-Flash |
| Input price (за M tokens) | $0.12 | $0.10 | $0.14 |
| Output price (за M tokens) | $0.24 | $0.20 | $0.28 |
| Prompt cache read (за M tokens) | $0.02 | $0.02 | $0.02 |
| Context window | 1M | 1M | 1M |
| Max output | 64K | 64K | 64K |
| License | Commercial API (Alibaba Cloud) | Open weights — MIT | Open weights — MIT |
Qwen 3.8 Flash на Alibaba Cloud (commercial API). GLM 5.3 Flash и DeepSeek V4 Flash — MIT open weights, self-hostable.
Сравнение бенчмарков
Flash tiers trade frontier reasoning за speed/cost — сравните честно перед production traffic.
| Метрика | Qwen 3.8 Flash | GLM 5.3 Flash | DeepSeek V4 Flash |
|---|---|---|---|
| SWE-bench Verified | 58.2% | 52.4% | 49.8% |
| HumanEval | 91.4% | 88.6% | 87.2% |
| Latency p50 (TTFT) | 180ms | 165ms | 140ms |
| Throughput | 420 tok/s | 480 tok/s | 510 tok/s |
SWE-bench и HumanEval vendor-reported на сентябрь 2026. Latency на regional edge — YMMV.
flash_models_page.premium_title
flash_models_page.premium_subtitle
Когда escalate на Claude Fable 5.1
Flash models — bulk routing, drafts, simple coding. Когда SWE-bench gaps matter — Fable 5.1 $10.00/$50.00/M даёт 87.7% SWE-bench (87.7%) vs ~50% на flash tiers. Hard 10–20% upstream.
Как выбрать
Default flash lane + escalate failures — не гоните всё по frontier pricing.
High-volume bulk routing
DeepSeek V4 Flash 510 tok/s, $0.14/$0.28 — classification, summarization, simple codegen at scale.
China-hosted commercial API
Qwen 3.8 Flash на Alibaba Cloud, RMB billing, mainland data residency. Когда нужен vendor SLA без self-hosting.
Self-hosting с open weights
GLM 5.3 Flash (MIT) и DeepSeek V4 Flash (MIT) на modest GPU fleets. GLM 5.3 Flash — cheapest list $0.10/$0.20.
Побеждает lowest list price
GLM 5.3 Flash $0.10/$0.20 undercuts Qwen и DeepSeek на input. Undifferentiated batch — start here.
Hard agentic coding needs escalation
Multi-file autonomous agents, complex refactors, frontier debugging — Fable 5.1, не flash tiers capped ~50–58% SWE-bench.
При 100M tokens/month разрыв $0.10 vs $10 input реален — как и cost failed agent retries.
Смоделируйте реальные costs
Сравните flash routing vs Fable 5.1 escalation с вашим token mix.