DeepSeek

DeepSeek-V4 Flash

The fast, low-cost V4 tier for high-volume tasks, available across inference gateways.

ActiveReleased April 24, 2026Available from 32 providers
Tool callingCan invoke external tools and APIs (function calling) mid-response.Structured outputCan return JSON or schema-constrained responses instead of free text.ReasoningUses extended thinking / chain-of-thought for harder multi-step problems.Effort controlLets you dial thinking depth (low / medium / high) to trade cost for quality.Fine-tuningSupports training a custom variant on your own data (vendor fine-tune API or self-hosted open weights).

Source: DeepSeek API models, DeepSeek models, Metadata verification, Field resolution verification

Price per 1M tokens
$0.14 in / $0.28 out
Context window131,072
Max output32,768
Cache read$0.028 / M
Batch discount-50%
Input
  • Text
Output
  • Text
Priced across32 providers

Pricing

Every billing leg from the live catalog, in USD per 1M tokens.

Billing legRateBatch
Input$0.14 / M$0.07 / M
Output$0.28 / M$0.14 / M
Cache read$0.028 / Mn/a

Cost calculator

Billing-grade math over the rate card above, including the thinking tokens most estimates miss.

Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.

Estimated monthly cost
$44.80
at current list prices
Fresh input$28.00
Cached input$0.00
Output incl. thinking$16.80
Track this automatically →

Compare providers

The same model, priced across every platform that serves it. Lowest combined in/out rate is flagged.

PlatformInput / MOutput / MCache read / MBatchAccess
Fireworks AI
$0.14$0.28$0.028✓Direct
StreamLakeLowest
$0.028$0.056$0.0056-via OpenRouter
DeepInfra
$0.09$0.18$0.018-via OpenRouter
Sail-research
$0.09$0.18$0.02-via OpenRouter
GMICloud
$0.091$0.182$0.0182-via OpenRouter
Venice
$0.0966$0.1925$0.0196-via OpenRouter
Akashml
$0.098$0.196$0.014-via OpenRouter
DigitalOcean
$0.098$0.196$0.0196-via OpenRouter
Baidu
$0.0983$0.1966$0.0197-via OpenRouter
Wafer
$0.1$0.25$0.05-via OpenRouter
Alibaba
$0.134$0.268$0.0268-via OpenRouter
SiliconFlow
$0.13$0.28$0.028-via OpenRouter
Morph
$0.139$0.278$0.07-via OpenRouter
Ambient
$0.14$0.28$0.028-via OpenRouter
AtlasCloud
$0.14$0.28$0.028-via OpenRouter
Coreweave
$0.14$0.28$0.07-via OpenRouter
Fireworks AI
$0.14$0.28$0.028-via OpenRouter
Ionstream
$0.14$0.28$0.07-via OpenRouter
Novita
$0.14$0.28$0.028-via OpenRouter
Parasail
$0.14$0.28$0.07-via OpenRouter
Wandb
$0.14$0.28$0.07-via OpenRouter
NextBit
$0.15$0.3$0.035-via OpenRouter
Phala
$0.2$0.4$0.07-via OpenRouter
Io-net
$0.239$0.379$0.099-via OpenRouter
Mancer-2
$0.175$0.5--via OpenRouter
Azure AI Foundry
$0.19$0.51$0.028-Direct
Azure AI Foundry · Data zone
$0.21$0.56$0.031-Direct
Azure OpenAI
$0.21$0.56$0.031-via OpenRouter
DeepSeek
$0.22$0.66$0.007-via OpenRouter
Openinference
$0.005$1.25$0.005-via OpenRouter
Relace
$0.0045$1.28$0.0045-via OpenRouter
Cloudflare
$0.44$1.32$0.014-via OpenRouter

Price history

Has this model gotten cheaper or more expensive?

EffectiveInput / OutputStatus
May 21, 2026 – July 10, 2026$0.14 / $0.28past
July 10, 2026$0.14 / $0.28current

More from DeepSeek

Other models from the same provider.

Stop estimating. Track it.

SuperPenguin meters every DeepSeek-V4 Flash call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.

Get started free

Rates are generated from SuperPenguin's live pricing catalog.