Global AI Token Plan comparison home

Deepseek Token API Pricing, Top-up, Deals & Official Entry

This DeepSeek API page compares official access, API key application, pricing and cache-hit deals rather than a monthly Coding Plan: under peak/off-peak pricing, V4-Flash costs ¥1.5 off-peak/¥3 peak input and ¥4.5/¥9 output per 1M tokens (0731 version) for high-volume batch and daily completion work, while V4-Pro costs ¥4.5/¥9 input and ¥13.5/¥27 output for complex reasoning, refactors and production agents; the experimental multimodal DeepSeek-V4-Flash-Vision is priced like Flash and recognizes images, screenshots and charts. Both support 1M context, up to 384K output, thinking/non-thinking modes, OpenAI / Anthropic compatibility, Tool Calls and usage limits, with cache-hit input dropping to ¥0.05–0.15 off-peak per 1M tokens, plus Claude Code, Cursor, Cline, Codex CLI and Volcengine / Tianyi Cloud / Qianfan bundle comparisons.

Token APIPay-as-you-go · 1M context

Deepseek Token API Latest Updates

Model

DeepSeek V4-Flash-Vision Exp Launches: Image-Recognizing Multimodal Model Priced Like V4-Flash

DeepSeek's experimental multimodal model deepseek-v4-flash-vision-exp went live on the API platform on 2026-08-21, adding image understanding on top of V4-Flash text capability. It handles JPEG/PNG/GIF/WebP, reads screenshot text, and analyzes charts, with vision-agent benchmarks leaping to near Opus-4.8. Billing is identical to V4-Flash: images are tokenized by size, capped at 384 tokens each, at ¥1.5 off-peak/¥3 peak input and ¥4.5/¥9 output per 1M tokens, concurrency 2500. This covers how to pass images, how billing works, official benchmarks, and the limits to watch.

Below is the complete pricing comparison for Deepseek Token API, covering 3 plan tiers, core model capabilities, quotas, and official entry. All data is sourced from the official website to help you decide whether it’s worth it.

Last updated

Deepseek Token API Price Comparison: Core Models

DeepSeek-V4-FlashDeepSeek-V4-ProDeepSeek-V4-Flash-Vision
DeepSeek-V4-Flash

Official fast API tier (deepseek-v4-flash): ¥1.5 off-peak/¥3 peak input, ¥4.5/¥9 output per million (cache miss), ¥0.05/¥0.10 cache-hit input; 1M context, concurrency 2500, thinking/non-thinking modes.

DeepSeek-V4-Pro

Official flagship API (deepseek-v4-pro): ¥4.5 off-peak/¥9 peak input, ¥13.5/¥27 output per million (cache miss), ¥0.15/¥0.30 cache-hit input; 1M context, up to 384K output, concurrency 500.

DeepSeek-V4-Flash-Vision

Official experimental multimodal API (deepseek-v4-flash-vision-exp): adds image understanding to V4-Flash text capability, JPEG/PNG/GIF/WebP, up to 600 images per request at 8192px max edge; billed like V4-Flash with images converted to tokens by size (384 max each), ¥1.5 off-peak/¥3 peak input (cache miss), ¥4.5/¥9 output per 1M tokens.

Deepseek Token API Price Comparison: Plans

DeepSeek-V4-Flash
Fast & cost-effective
Input
¥1.5–3
Output
¥4.5–9
Input ¥1.5 off-peak/¥3 peak · output ¥4.5/¥9 · cache-hit input ¥0.05/¥0.10 · concurrency 2500
Usage
Model id deepseek-v4-flash (current version -0731): ¥1.5 off-peak/¥3 peak input (cache miss), ¥4.5/¥9 output—suited to high-frequency batch work, daily completion, and many light requests.
Models
1M context, up to 384K output, concurrency cap 2500; supports JSON Output, Tool Calls, Responses API, chat prefix completion—FIM completion only in non-thinking mode.
Highlights
Thinking/non-thinking modes (thinking default); cache-hit input as low as ¥0.05 off-peak/M (¥0.10 peak)—build caching into architecture for repeated-context workloads.
Suited to self-built backends, script automation, agent pipelines, log analysis, and users who need tight per-call cost control.
Best for
High-volume batch callers, self-built backend developers, and daily completion users optimizing for cost
DeepSeek-V4-Pro
Flagship reasoningRecommended
Input
¥4.5–9
Output
¥13.5–27
Input ¥4.5 off-peak/¥9 peak · output ¥13.5/¥27 · cache-hit input ¥0.15/¥0.30 · concurrency 500
Usage
Model id deepseek-v4-pro (current version -0813): ¥4.5 off-peak/¥9 peak input (cache miss), ¥13.5/¥27 output—costs more than Flash, suited to complex reasoning, deep refactors, and long agent chains.
Models
1M context, up to 384K output, concurrency cap 500; same thinking/non-thinking switch, JSON Output, Tool Calls, and FIM completion (non-thinking only).
Highlights
Cache-hit input ¥0.15 off-peak/M (¥0.30 peak) still lowers cost for repeated context—reserve Pro for requests that truly need stronger capability.
Suited to complex reasoning and coding tasks, professional developers, and production systems treating DeepSeek as a core model base.
Best for
Complex reasoning and coding users, professional developers, and builders of production-grade agent systems
DeepSeek-V4-Flash-Vision
Experimental multimodal
Input
¥1.5–3
Output
¥4.5–9
Input ¥1.5 off-peak/¥3 peak · output ¥4.5/¥9 · cache-hit input ¥0.05/¥0.10 · concurrency 2500
Usage
Model id deepseek-v4-flash-vision-exp (experimental): adds image understanding on top of V4-Flash text capability, supports JPEG/PNG/GIF/WebP, reads screenshot text, analyzes charts, and drives Agent tools.
Models
Billing is identical to V4-Flash: ¥1.5 off-peak/¥3 peak input (cache miss), ¥4.5/¥9 output per 1M tokens; images convert to tokens by size, capped at 384 per image, keeping cost predictable.
Highlights
Up to 600 images per request, 8192px max edge, accepted via Base64 / URL / Files API; 1M context, 384K max output, concurrency 2500, compatible with Chat Completions / Responses / Anthropic formats.
Experimental: currently only this model accepts images, no FIM completion, images only in user messages; passing an image to regular V4-Flash / V4-Pro errors.
Best for
Screenshot understanding, mixed text-image document processing, chart analysis, and vision-capable agents

Deepseek Token API Price Comparison: Notes

  • No official monthly coding plan—charges = token usage × unit price, deducted from top-up or gift balance; gift balance is used first when both exist.
  • Pricing page lists deepseek-v4-flash, deepseek-v4-pro and the experimental vision model deepseek-v4-flash-vision-exp; deepseek-chat / deepseek-reasoner deprecate 2026-07-24 23:59 CST, mapping to v4-flash non-thinking and thinking modes.
  • Concurrency per account: v4-flash 2500, v4-pro 500, v4-flash-vision-exp 2500; HTTP 429 when exceeded. DeepSeek also appears in Volcengine, CtCloud, Qianfan coding plans via plan quota—not official balance.
  • Peak/off-peak pricing: since 2026-08-17 00:00 CST, off-peak prices are half the peak rate, with peak hours 9:00–12:00 and 14:00–18:00 CST (all other times are off-peak). Prices on this page show off-peak–peak ranges; actual charges follow the official billing page.
  • The experimental vision model deepseek-v4-flash-vision-exp launched 2026-08-21: it's currently the only model accepting images, has no FIM completion yet, and images may only appear in user messages; images convert to tokens by size (384 max), up to 600 per request at 8192px max edge, and can be reused via file_id through the free Files API.

Deepseek Token API Price Comparison: Tools & Integration

OpenAI-compatible APIAnthropic-compatible APIClaude CodeCursorClineCodex CLI

Deepseek Token API Pricing, Top-up, Deals, Quota, Usage, Setup & Updates

AnalysisPublished Updated

Where to Find DeepSeek Harness Plugins? DSH-Plugin Hub Indexes 3,000+ dsh-plugin Plugins and Updates Daily

DeepSeek Harness pluginsDSH plugindsh-pluginplugin directoryplugin installationDSH-Plugin Hub

Where to find and install DeepSeek Harness plugins? The community-maintained DSH-Plugin Hub plugin directory indexes 3,000+ dsh-plugin entries across 11 categories (Tools & Capabilities, Skills & Agents, Models & Reasoning and more), links every plugin back to its GitHub source, updates daily and marks verified compatibility — far more efficient than raw-searching dsh-plugin on GitHub. This article explains how the plugin directory works and gives the one-line dsh plugin add install command so you can find and install a solid DeepSeek Harness plugin in minutes.

Page published

Deepseek Token API Pricing, Top-up, Deals & FAQ

AI Token Plan

Global AI Token Comparison

AI Token Plan is not just a collection of vendor links. It places 41+ domestic and international AI platforms into one comparison framework, covering 395+ model entries plus common plan types such as Token Plans, Coding Plans, IDE Tools and LLM APIs.

When you need to compare AI subscription pricing, coding plan quota, API usage cost or official deal entry points, AI Token Plan brings official prices, plan tiers, usage rules, model capabilities and tool integrations into one place, reducing the need to check multiple vendor sites manually.

Data is continuously organized as vendor pricing pages, plan pages and product documentation change, making the homepage a pricing comparison entry point while detail pages explain whether each vendor plan fits individual developers, team purchasing or long-term API usage.

Prices come from each platform's official site and may change at any time; the official price prevails.

© 2026 AI Token Plan · All rights reserved · First published June 18, 2026 · 64 days running · Sitemap