DeepSeek
DeepSeek V4-Flash
DeepSeek · open-weight. Batchable work you can schedule outside DeepSeek's peak hours, where the rate halves to one of the lowest hosted prices among models tracked by benchr — not a market-wide guarantee.
The published record
Every number below was read from the provider's own documentation on the date shown. benchr does not restate a figure it has not seen published.
- API identifier
deepseek-v4-flash- Context window
- 1MtokensSource
- Maximum output
- 384KtokensSource
- Input
- $0.44per 1M tokensSource
- Output
- $1.32per 1M tokensSource
- Released
- April 24, 2026Source
- License
- MIT
Record verified August 28, 2026
What the record says
UPDATED 2026-08-28: DeepSeek raised API prices and introduced peak/off-peak billing effective 16:00 UTC on August 16, 2026, announced alongside the V4 lineup release. Cache-miss input went from a flat $0.14 to $0.44 peak / $0.22 off-peak per 1M tokens, output from $0.28 to $1.32 / $0.66, and cache-hit input from $0.0028 to $0.014 / $0.007. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday, so roughly four fifths of the week bills at the off-peak rate; that share is a benchr calculation from the published hours, not a provider figure. SWE-bench Verified 79.0% is DeepSeek's provider-reported model-card figure, not an independently reproduced result. Open weights, MIT.
What has moved
Entries in the change ledger that affect this model, newest first.
Price changed
Cache-miss input rose from a flat $0.14 to $0.44 peak / $0.22 off-peak per 1M tokens, output from $0.28 to $1.32 / $0.66, and cache-hit input from $0.0028 to $0.014 / $0.007. DeepSeek replaced flat API pricing with peak/off-peak billing at 16:00 UTC on August 16, 2026. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday; all other hours bill at half the peak rate. The pricing figures here are the peak (list) rates. Source: api-docs.deepseek.com/quick_start/pricing (read 2026-08-28).
Released
DeepSeek V4-Flash launched April 24, 2026. MIT license. Cheapest commercial API in this set. Source: api-docs.deepseek.com
Which one to use
The rest of the family, with the two numbers that usually decide it. The current model is marked.
| Models | Input | Output | Context window |
|---|---|---|---|
| DeepSeek V4-Flash This page | $0.44 | $1.32 | 1M |
| DeepSeek V4-Pro | $1.32 | $3.96 | 1M |
benchr's read
Worth it for
- Batchable work you can schedule outside DeepSeek's peak hours, where the rate halves to one of the lowest hosted prices among models tracked by benchr — not a market-wide guarantee
- Self-hosted at low cost
- Volume coding tasks
Look elsewhere if
- You need frontier-tier reasoning depth
- Vision-first workflows
- Your traffic is fixed to European or Asian business hours, which sit inside the peak window
Prices, limits, and identifiers above are provider facts. This section is benchr's judgement about them.
What benchr has written about it
Pieces that name this model, newest first. The ones written about this model come before the ones that mention it in passing.
What is not on this page
Stated rather than filled in.
- No first-token or tokens-per-second figure is recorded. benchr has not measured it and the provider does not publish one.