Deepseek Token API Pricing, Top-up, Deals & Official Entry
This DeepSeek API page compares official access, API key application, pricing and cache-hit deals rather than a monthly Coding Plan: under peak/off-peak pricing, V4-Flash costs ¥1.5 off-peak/¥3 peak input and ¥4.5/¥9 output per 1M tokens (0731 version) for high-volume batch and daily completion work, while V4-Pro costs ¥4.5/¥9 input and ¥13.5/¥27 output for complex reasoning, refactors and production agents; the experimental multimodal DeepSeek-V4-Flash-Vision is priced like Flash and recognizes images, screenshots and charts. Both support 1M context, up to 384K output, thinking/non-thinking modes, OpenAI / Anthropic compatibility, Tool Calls and usage limits, with cache-hit input dropping to ¥0.05–0.15 off-peak per 1M tokens, plus Claude Code, Cursor, Cline, Codex CLI and Volcengine / Tianyi Cloud / Qianfan bundle comparisons.
Below is the complete pricing comparison for Deepseek Token API, covering 3 plan tiers, core model capabilities, quotas, and official entry. All data is sourced from the official website to help you decide whether it’s worth it.
Last updated:
Deepseek Token API Price Comparison: Core Models
Official fast API tier (deepseek-v4-flash): ¥1.5 off-peak/¥3 peak input, ¥4.5/¥9 output per million (cache miss), ¥0.05/¥0.10 cache-hit input; 1M context, concurrency 2500, thinking/non-thinking modes.
Official flagship API (deepseek-v4-pro): ¥4.5 off-peak/¥9 peak input, ¥13.5/¥27 output per million (cache miss), ¥0.15/¥0.30 cache-hit input; 1M context, up to 384K output, concurrency 500.
Official experimental multimodal API (deepseek-v4-flash-vision-exp): adds image understanding to V4-Flash text capability, JPEG/PNG/GIF/WebP, up to 600 images per request at 8192px max edge; billed like V4-Flash with images converted to tokens by size (384 max each), ¥1.5 off-peak/¥3 peak input (cache miss), ¥4.5/¥9 output per 1M tokens.
Deepseek Token API Price Comparison: Plans
Suited to self-built backends, script automation, agent pipelines, log analysis, and users who need tight per-call cost control.
Suited to complex reasoning and coding tasks, professional developers, and production systems treating DeepSeek as a core model base.
Experimental: currently only this model accepts images, no FIM completion, images only in user messages; passing an image to regular V4-Flash / V4-Pro errors.
Deepseek Token API Price Comparison: Notes
- No official monthly coding plan—charges = token usage × unit price, deducted from top-up or gift balance; gift balance is used first when both exist.
- Pricing page lists deepseek-v4-flash, deepseek-v4-pro and the experimental vision model deepseek-v4-flash-vision-exp; deepseek-chat / deepseek-reasoner deprecate 2026-07-24 23:59 CST, mapping to v4-flash non-thinking and thinking modes.
- Concurrency per account: v4-flash 2500, v4-pro 500, v4-flash-vision-exp 2500; HTTP 429 when exceeded. DeepSeek also appears in Volcengine, CtCloud, Qianfan coding plans via plan quota—not official balance.
- Peak/off-peak pricing: since 2026-08-17 00:00 CST, off-peak prices are half the peak rate, with peak hours 9:00–12:00 and 14:00–18:00 CST (all other times are off-peak). Prices on this page show off-peak–peak ranges; actual charges follow the official billing page.
- The experimental vision model deepseek-v4-flash-vision-exp launched 2026-08-21: it's currently the only model accepting images, has no FIM completion yet, and images may only appear in user messages; images convert to tokens by size (384 max), up to 600 per request at 8192px max edge, and can be reused via file_id through the free Files API.
Deepseek Token API Price Comparison: Tools & Integration
Deepseek Token API Pricing, Top-up, Deals, Quota, Usage, Setup & Updates
Page published: