DeepSeek V4.1 Flash pricing
Budget LLM from DeepSeek —$0.15 input and$0.60 output per 1M tokens.
GABudgetOff-peak rate shown
- Input /1M
- $0.15
- Output /1M
- $0.60
- Cached /1M
- $0.003
- Blended
- $0.375
- Context
- 1M
- Max output
- 384K
API model id: deepseek-flash
Full price matrix
| Rate | Input /1M | Output /1M | Notes |
|---|---|---|---|
| Standard | $0.15 | $0.60 | |
| Cached input (read) | $0.003 | — | Prompt-cache hit |
| Peak (×2) | $0.30 | $1.20 | During peak hours |
List prices, USD, directional. Rates are provider list prices per 1M tokens and are meant for comparison, not billing. Batch rates shown as 50% off are derived where a provider offers batch but does not publish a separate figure. Preview, promo, intro, peak/off-peak, long-context, and third-party-host prices are labeled where they apply. Token counts vary by tokenizer, so per-token price is not always a like-for-like cost. Always confirm with the provider before relying on a number. Prices as of 2026-09-29.
Source: https://api-docs.deepseek.com/quick_start/pricing
Pricing notes
- Off-peak / peak pricing: Launched 2026-09-10 under the stable API id deepseek-flash, replacing V4 Flash. Off-peak rate shown; peak rate is 2× ($0.30 / $1.20, cached $0.006). Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday, excluding Chinese public holidays. See off-peak pricing.