DeepSeek

DeepSeek-V4.1-Flash

DeepSeek · open-weight. Replacing the retired DeepSeek V4-Flash on the same API, at a lower listed rate.

The published record

Every number below was read from the provider's own documentation on the date shown. benchr does not restate a figure it has not seen published.

API identifier
deepseek-flash
Context window
1MtokensSource
Maximum output
384KtokensSource
Input
$0.3per 1M tokensSource
Output
$1.20per 1M tokensSource
Released
September 10, 2026Source
License
Not stated

Availability Available on the DeepSeek API since September 10, 2026, replacing DeepSeek-V4-Flash and V4-Flash-Vision-Exp.

Record verified September 11, 2026

What the record says

DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026 under the API ID deepseek-flash, describing it as a 552B-parameter mixture-of-experts model with 8B active parameters for input and 16B for output. The pricing page lists a 1M context window, 384K maximum output and 2,500 concurrent requests. It replaced DeepSeek-V4-Flash and V4-Flash-Vision-Exp, whose legacy IDs temporarily route to it. The release note does not state a license or link open weights, so license is null rather than copied from V4-Flash's MIT weights, and the active-parameter field is null because DeepSeek gives two figures rather than one.

Which one to use

The rest of the family, with the two numbers that usually decide it. The current model is marked.

ModelsInput OutputContext window
DeepSeek-V4.1-Flash This page$0.3$1.201M
DeepSeek V4-Pro$1.32$3.961M

benchr's read

Worth it for

  • Replacing the retired DeepSeek V4-Flash on the same API, at a lower listed rate
  • Batchable work scheduled outside DeepSeek's peak hours, where every rate halves to among the lowest hosted rates among models tracked by benchr - not a market-wide guarantee
  • Long-context work that needs a 1M-token window and up to 384K output tokens

Look elsewhere if

  • You need published benchmark figures - DeepSeek released none that benchr records
  • You need a stated license or open weights for this release - DeepSeek published neither
  • Your traffic is fixed to European or Asian business hours, which sit inside the peak window

Prices, limits, and identifiers above are provider facts. This section is benchr's judgement about them.

What benchr has written about it

Pieces that name this model, newest first. The ones written about this model come before the ones that mention it in passing.

What is not on this page

Stated rather than filled in.

  • No first-token or tokens-per-second figure is recorded. benchr has not measured it and the provider does not publish one.