価格 2026年9月時点で照合

3 flash モデルとも vendor リスト価格、2026年9月。ベンチは特記なき限り vendor 自行報告。

コスト計算機

Flash 三者比較 · 2026年9月

Qwen 3.8 Flash vs GLM 5.3 Flash vs DeepSeek V4 Flash

2026年9月時点で最速の中国 lab commodity モデル3つ — 価格、スループット、coding 品質、cheap flash routing が frontier に勝つ条件。

TL;DR

GLM 5.3 Flash が最安($0.10/$0.20)MIT open weights。Qwen 3.8 Flash が flash tier で SWE-bench 58.2% 先行。DeepSeek V4 Flash が raw latency 140ms p50。hard agent work は Fable 5.1 へ。

スペック概要

100万トークンあたり USD、2026年9月 vendor リスト。

指標Qwen 3.8 FlashGLM 5.3 FlashDeepSeek V4 Flash
モデル IDQwen-3.8-FlashGLM-5.3-FlashDeepSeek-V4-Flash
入力価格(100万トークン)$0.12$0.10$0.14
出力価格(100万トークン)$0.24$0.20$0.28
Prompt cache read(100万トークン)$0.02$0.02$0.02
コンテキスト1M1M1M
最大出力64K64K64K
ライセンスCommercial API(Alibaba Cloud)Open weights — MITOpen weights — MIT

Qwen 3.8 Flash は Alibaba Cloud(commercial API)。GLM 5.3 Flash と DeepSeek V4 Flash は MIT open weights、self-host 可。

ベンチマーク比較

Flash tier は frontier reasoning を speed/cost と trade — production 前に honest 比較。

指標Qwen 3.8 FlashGLM 5.3 FlashDeepSeek V4 Flash
SWE-bench Verified58.2%52.4%49.8%
HumanEval91.4%88.6%87.2%
Latency p50(TTFT)180ms165ms140ms
スループット420 tok/s480 tok/s510 tok/s

SWE-bench/HumanEval は vendor 自行報告、2026年9月。latency は regional edge 測定 — 環境で変動。

flash_models_page.premium_title

flash_models_page.premium_subtitle

Claude Fable 5.1 に escalate するタイミング

Flash は bulk routing、draft、簡単 coding。SWE-bench gap が効くとき — Fable 5.1 $10.00/$50.00/M で 87.7% SWE-bench(87.7%)、flash tier は約50%。hard 10–20% を upstream。

選び方

デフォルト flash lane を選び、失敗で escalate — 全部 frontier 価格で回さない。

  • 高 volume bulk routing

    DeepSeek V4 Flash 510 tok/s、$0.14/$0.28 — 分類、要約、簡単 codegen の scale。

  • 中国 hosted commercial API

    Qwen 3.8 Flash on Alibaba Cloud、RMB billing、mainland data residency。SLA 必要で self-host しない場合。

  • open weights self-host

    GLM 5.3 Flash(MIT)と DeepSeek V4 Flash(MIT)は modest GPU fleet で可。GLM 5.3 Flash が最安 $0.10/$0.20。

  • 最低リスト価格優先

    GLM 5.3 Flash $0.10/$0.20 が Qwen/DeepSeek 入力より安い。無差別 batch はここから。

  • hard agentic coding は escalate

    multi-file 自律 agent、complex refactor、frontier debug は Fable 5.1 — SWE-bench 50–58% cap の flash tier ではない。

月1億 token なら $0.10 と $10 入力差は real — agent 失敗リトライも real。

実コストモデル

flash routing vs Fable 5.1 escalate を自社 token mix で比較。

FAQ