Цены сверены на сентябрь 2026

Все три flash — vendor list rates на сентябрь 2026. Бенчмарки vendor-reported unless noted.

Калькулятор стоимости

Flash models three-way · сентябрь 2026

Qwen 3.8 Flash vs GLM 5.3 Flash vs DeepSeek V4 Flash

Три fastest Chinese-lab commodity models на сентябрь 2026 — цена, throughput, coding quality, когда cheap flash routing beats frontier.

Кратко

GLM 5.3 Flash cheapest ($0.10/$0.20) с MIT open weights. Qwen 3.8 Flash лидирует SWE-bench среди flash tiers (58.2%). DeepSeek V4 Flash — raw latency 140ms p50. Hard agent work — escalate на Fable 5.1.

Спеки одним взглядом

USD за M tokens, vendor list на сентябрь 2026.

МетрикаQwen 3.8 FlashGLM 5.3 FlashDeepSeek V4 Flash
Model IDQwen-3.8-FlashGLM-5.3-FlashDeepSeek-V4-Flash
Input price (за M tokens)$0.12$0.10$0.14
Output price (за M tokens)$0.24$0.20$0.28
Prompt cache read (за M tokens)$0.02$0.02$0.02
Context window1M1M1M
Max output64K64K64K
LicenseCommercial API (Alibaba Cloud)Open weights — MITOpen weights — MIT

Qwen 3.8 Flash на Alibaba Cloud (commercial API). GLM 5.3 Flash и DeepSeek V4 Flash — MIT open weights, self-hostable.

Сравнение бенчмарков

Flash tiers trade frontier reasoning за speed/cost — сравните честно перед production traffic.

МетрикаQwen 3.8 FlashGLM 5.3 FlashDeepSeek V4 Flash
SWE-bench Verified58.2%52.4%49.8%
HumanEval91.4%88.6%87.2%
Latency p50 (TTFT)180ms165ms140ms
Throughput420 tok/s480 tok/s510 tok/s

SWE-bench и HumanEval vendor-reported на сентябрь 2026. Latency на regional edge — YMMV.

flash_models_page.premium_title

flash_models_page.premium_subtitle

Когда escalate на Claude Fable 5.1

Flash models — bulk routing, drafts, simple coding. Когда SWE-bench gaps matter — Fable 5.1 $10.00/$50.00/M даёт 87.7% SWE-bench (87.7%) vs ~50% на flash tiers. Hard 10–20% upstream.

Как выбрать

Default flash lane + escalate failures — не гоните всё по frontier pricing.

  • High-volume bulk routing

    DeepSeek V4 Flash 510 tok/s, $0.14/$0.28 — classification, summarization, simple codegen at scale.

  • China-hosted commercial API

    Qwen 3.8 Flash на Alibaba Cloud, RMB billing, mainland data residency. Когда нужен vendor SLA без self-hosting.

  • Self-hosting с open weights

    GLM 5.3 Flash (MIT) и DeepSeek V4 Flash (MIT) на modest GPU fleets. GLM 5.3 Flash — cheapest list $0.10/$0.20.

  • Побеждает lowest list price

    GLM 5.3 Flash $0.10/$0.20 undercuts Qwen и DeepSeek на input. Undifferentiated batch — start here.

  • Hard agentic coding needs escalation

    Multi-file autonomous agents, complex refactors, frontier debugging — Fable 5.1, не flash tiers capped ~50–58% SWE-bench.

При 100M tokens/month разрыв $0.10 vs $10 input реален — как и cost failed agent retries.

Смоделируйте реальные costs

Сравните flash routing vs Fable 5.1 escalation с вашим token mix.

FAQ