Verified 2026-10-07 · sourced from DeepSeek
DeepSeek V4.1 Flash Token Calculator, Pricing & 100K/1M Cost
Check DeepSeek DeepSeek V4.1 Flash pricing, estimate 100K and 1M token cost, and size a real API budget before you send a single request. Standard pricing is $0.30 per million input tokens and $1.20 per million output tokens with a 1M token context window.
Quick answer: DeepSeek V4.1 Flash pricing per 1M tokens is $0.30 input and $1.20 output. Context window: 1,000,000 tokens · Cached input: $0.006 / 1M.
Best for searches like DeepSeek V4.1 Flash token calculator, DeepSeek V4.1 Flash pricing, DeepSeek V4.1 Flash 100K tokens cost, DeepSeek V4.1 Flash 1M token cost.
Reference rates: Peak rates · select off-peak when applicable
USD per million text tokens. Verified 2026-10-07 · Official source
| Mode / prompt size | Input | Cache read | Cache write | Output |
|---|---|---|---|---|
| Peak | $0.3 | $0.006 | Not estimated | $1.2 |
| Off-peak | $0.15 | $0.003 | Not estimated | $0.6 |
1,000,000 context tokens · max output 384,000. Local token counts are estimates; use the provider’s usage counts for an invoice estimate.
Direct DeepSeek API. Peak: Monday–Friday 01:00–04:00 and 06:00–10:00 UTC, excluding Chinese public holidays. All other times are off-peak. Choose the applicable mode explicitly; no holiday calendar is inferred.
Cache hit and cache miss are separate input rates. Batch and cache-write prices have not been verified and are not estimated. Text token counts are approximate; provider-reported usage is authoritative.
Pick the route that matches what you searched for
Some visitors want a fast DeepSeek V4.1 Flash API cost estimate, others want a direct 100K or 1M token budget, and some are already comparing alternatives. These shortcuts remove the extra click.
Estimate a single request or prompt budget right now.
Jump straight to the most common budgeting checkpoint.
Use this when you are sizing production traffic or a monthly plan.
Open the closest head-to-head comparison instead of researching from scratch.
Context window
1,000,000 tokens
Input price
$0.30 / 1M
Output price
$1.20 / 1M
Cached input
$0.006 / 1M
Pricing modes and thresholds
Usage scenarios
Compare standard and cached pricing (where available) across common workloads.
| Scenario | Tokens in | Tokens out | Total tokens | Standard cost | Cached cost |
|---|---|---|---|---|---|
Quick chat reply Single user question with a short assistant answer | 650 | 220 | 870 | $0.0005 | $0.0003 |
Coding assistant session Multi-turn pair programming exchange (≈6 turns) | 2,600 | 1,400 | 4,000 | $0.0025 | $0.0017 |
Knowledge base response Retrieval-augmented answer referencing multiple passages | 12,000 | 3,000 | 15,000 | $0.0072 | $0.0037 |
Near-max context run Large document processing approaching the 1M token limit | 880,000 | 120,000 | 1,000,000 | $0.408 | $0.149 |
Daily & monthly budgeting
Translate usage into predictable operating expenses across popular deployment sizes.
| Profile | Requests/day | Tokens/day | Daily cost | Monthly cost | Cached daily | Cached monthly |
|---|---|---|---|---|---|---|
| Team pilot | 25 | 75,000 | $0.0450 | $1.35 | $0.0303 | $0.909 |
| Product launch | 100 | 500,000 | $0.285 | $8.55 | $0.182 | $5.46 |
| Enterprise scale | 500 | 3,000,000 | $1.80 | $54.00 | $1.21 | $36.36 |
Pricing notes
- Direct DeepSeek API. Peak: Monday–Friday 01:00–04:00 and 06:00–10:00 UTC, excluding Chinese public holidays. All other times are off-peak. Choose the applicable mode explicitly; no holiday calendar is inferred.
- Cache hit and cache miss are separate input rates. Batch and cache-write prices have not been verified and are not estimated. Text token counts are approximate; provider-reported usage is authoritative.
Frequently asked questions
How much does DeepSeek V4.1 Flash cost per 1,000 tokens?
At the published rates of $0.30 per million input tokens and $1.20 per million output tokens, a typical 1,000 token request (≈70% input, 30% output) costs about $0.0006.
Does DeepSeek V4.1 Flash offer cached input discounts?
DeepSeek V4.1 Flash drops input costs to $0.006 per million cached tokens. Using cached contexts, that same 1,000 token call totals $0.0004, a significant saving for chatbots and RAG systems.
What is the context window for DeepSeek V4.1 Flash?
DeepSeek V4.1 Flash supports up to 1,000,000 tokens (1M), allowing large prompts and retrieval-augmented payloads in a single call.
How fresh is the DeepSeek V4.1 Flash pricing data?
Pricing is sourced from https://api-docs.deepseek.com/quick_start/pricing/ and was last verified on 2026-10-07. The calculator updates automatically when models.json is refreshed.