DeepSeek V4 Flash
Summary
DeepSeek V4 Flash is an open-weight language model from DeepSeek, released on 24 Apr 2026. It accepts text and generates text, with a 1M-token context window and up to 16K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.09 per million input tokens and $0.18 per million output tokens (DeepInfra), across 7 providers we track.
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Last verified 4 Oct 2026. Every figure on this page links to its source.
Specifications
| Developer | DeepSeek |
|---|---|
| Release date | 24 Apr 2026 [models.dev] |
| Status | Generally available [models.dev] |
| Weights | Open weights [Hugging Face] |
| License | mit [Hugging Face] |
| Parameters | 291B [Hugging Face] |
| Context window | 1M (1,048,576 tokens) [models.dev] |
| Max output | 16K tokens [models.dev] |
| Knowledge cutoff | May 2025 [models.dev] |
| Input | Text [models.dev] |
| Output | Text [models.dev] |
| Reasoning mode | Yes [models.dev] |
| Tool calling | Yes [models.dev] |
| Structured output | Yes [models.dev] |
| Hugging Face | deepseek-ai/DeepSeek-V4-Flash [Hugging Face] |
DeepSeek V4 Flash API pricing by provider
USD per million tokens. Sorted by input price.
| Provider | Input | Output | Cached input | Context | Source |
|---|---|---|---|---|---|
| OpenRouter deepseek/deepseek-v4-flash | $0.02 | $1.28 | $0.02 | 1M | OpenRouter, 4 Oct 2026 |
| DeepInfra deepseek-ai/DeepSeek-V4-Flash | $0.09 | $0.18 | $0.02 | 1M | models.dev, 4 Oct 2026 |
| SiliconFlow deepseek-ai/DeepSeek-V4-Flash | $0.13 | $0.28 | $0.03 | 1M | models.dev, 4 Oct 2026 |
| Vercel AI Gateway deepseek/deepseek-v4-flash | $0.13 | $0.26 | $0.03 | 1M | models.dev, 4 Oct 2026 |
| Novita AI deepseek/deepseek-v4-flash | $0.14 | $0.28 | $0.03 | 1M | models.dev, 4 Oct 2026 |
| Hugging Face Inference Providers deepseek-ai/DeepSeek-V4-Flash | $0.14 | $0.28 | – | 1M | models.dev, 4 Oct 2026 |
| Azure AI Foundry deepseek-v4-flash | $0.19 | $0.51 | – | 1M | models.dev, 4 Oct 2026 |
Frequently asked questions
What is the context window of DeepSeek V4 Flash?
DeepSeek V4 Flash has a context window of 1,048,576 tokens (1M) and can generate up to 16,384 tokens in a single response, according to models.dev.
How much does the DeepSeek V4 Flash API cost?
DeepSeek V4 Flash is listed from $0.09 per million input tokens and $0.18 per million output tokens (DeepInfra). Prices last verified 4 Oct 2026.
Is DeepSeek V4 Flash open source?
DeepSeek V4 Flash is an open-weight model: its weights are published on Hugging Face as deepseek-ai/DeepSeek-V4-Flash under the mit license. Check the license terms before commercial use.
When was DeepSeek V4 Flash released?
DeepSeek released DeepSeek V4 Flash on 24 Apr 2026, according to models.dev.
Which providers offer DeepSeek V4 Flash?
We track DeepSeek V4 Flash on 7 providers: OpenRouter, DeepInfra, SiliconFlow, Vercel AI Gateway, Novita AI, Hugging Face Inference Providers, Azure AI Foundry.
Change history
No changes detected since we started tracking this model. We re-check every source daily.
Sources
- https://openrouter.ai/deepseek/deepseek-v4-flash
- https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/deepseek-ai/DeepSeek-V4-Flash.toml
- https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash
- https://github.com/sst/models.dev/blob/dev/providers/siliconflow/models/deepseek-ai/DeepSeek-V4-Flash.toml
- https://github.com/sst/models.dev/blob/dev/providers/vercel/models/deepseek/deepseek-v4-flash.toml
- https://github.com/sst/models.dev/blob/dev/providers/novita-ai/models/deepseek/deepseek-v4-flash.toml
- https://github.com/sst/models.dev/blob/dev/providers/huggingface/models/deepseek-ai/DeepSeek-V4-Flash.toml
- https://github.com/sst/models.dev/blob/dev/providers/azure/models/deepseek-v4-flash.toml