DeepSeek
DeepSeek V4 Flash
Fastest, cheapest 1M-context model on the gateway. The new default value pick.
ValueFastLong context
$ / 1M input
$0.140
What you pay per million prompt tokens.
$ / 1M output
$0.280
What you pay per million completion tokens.
Blended (70 / 30)
$0.182
Typical chat workload mix.
Ratings & benchmarks
Snapshot 2026-04-28Overall
4.6 / 5
Composite of intelligence + reliability.
Output speed
130 t/s
Output tokens per second under typical load.
Time to first token
600 ms
Lower is better. Reasoning models naturally stretch this.
Value
5.0 / 5
Intelligence per dollar at typical mix.
Context window
1,048,576 tokens