Deepseek Token API Pricing, Top-up, Deals & Updates | Article

Deepseek Token API pricing. DeepSeek-V4-Flash, DeepSeek-V4-Pro and 1 more models. OpenAI-compatible API, Anthropic-compatible API, Claude Code and 3 more tools integrations. Includes cache tiers and rate limits. API key setup, usage quotas, and official entry—is it worth it?

41
Platforms
395+
Models
140+
IDE Tools
3
Deepseek Token API models
¥1.5–3 / ¥4.5–9
Deepseek Token API entry price
View
Vendor details

DeepSeek API Latest Updates

DeepSeek API Price Hike Is Live: Peak/Off-Peak Pricing, V4-Flash vs V4-Pro New Rates, Markup and How to Save

Published Updated
PricingDeepSeekAPIPrice IncreasePeak/Off-Peak PricingPricing ComparisonV4-FlashV4-Pro
Summary

DeepSeek API officially switched to peak/off-peak pricing on August 17, 2026, and prices went up across the board. During peak hours (9:00–12:00 and 14:00–18:00 CST), v4-flash input rises to ¥3 per 1M tokens and output to ¥9, while v4-pro hits ¥9 input and ¥27 output; off-peak is halved. Cache-hit input also climbed from ¥0.02–0.025 to ¥0.05–0.30. This guide compares old vs new rates, breaks down the markup, explains top-up, quota, rate limits and API key setup, then covers off-peak scheduling, caching and third-party alternatives like Volcengine so you can decide whether DeepSeek is still worth it.

This isn't a small bump — DeepSeek lifted the whole pricing table a notch. The official pricing page posted a notice on August 15, 2026: starting 2026-08-17 00:00 CST, billing switches to peak/off-peak pricing, with off-peak at half the peak rate (source). Peak hours are 9:00–12:00 and 14:00–18:00 CST — right in the middle of your coding and batch hours.

What the New Prices Look Like

Peak rates start at 3× and output goes straight to 4.5×. Peak/off-peak pricing splits the day into two tiers — peak at full price, off-peak at half. All new prices below are in CNY (¥), matching DeepSeek's official announcement.

Peak hours (9:00–12:00, 14:00–18:00)

Item V4-Flash V4-Pro
Cache-miss input ¥3 / 1M tokens ¥9 / 1M tokens
Output ¥9 / 1M tokens ¥27 / 1M tokens
Cache-hit input ¥0.10 / 1M tokens ¥0.30 / 1M tokens

Off-peak hours (all other times)

Item V4-Flash V4-Pro
Cache-miss input ¥1.5 / 1M tokens ¥4.5 / 1M tokens
Output ¥4.5 / 1M tokens ¥13.5 / 1M tokens
Cache-hit input ¥0.05 / 1M tokens ¥0.15 / 1M tokens

Both models keep the same specs: 1M context, up to 384K output, thinking/non-thinking modes, OpenAI and Anthropic format support. Concurrency is unchanged at 2500 for v4-flash and 500 for v4-pro (source).

How Big the Markup Is

At peak, v4-flash input is 3× and output 4.5×, while v4-pro cache-hit input jumped 12×. Side by side, the change looks like this (all CNY):

Item Model Before Off-peak Peak Peak markup
Cache-miss input V4-Flash ¥1 ¥1.5 ¥3
Output V4-Flash ¥2 ¥4.5 ¥9 4.5×
Cache-hit input V4-Flash ¥0.02 ¥0.05 ¥0.10
Cache-miss input V4-Pro ¥3 ¥4.5 ¥9
Output V4-Pro ¥6 ¥13.5 ¥27 4.5×
Cache-hit input V4-Pro ¥0.025 ¥0.15 ¥0.30 12×

The ugliest line is v4-pro's cache-hit input — up from ¥0.025 to ¥0.30 at peak, a 12× jump. Workloads that used to live on cache discounts need a fresh cost model.

Top-up, Billing, Quota and Rate Limits Are Unchanged

Top-up and billing logic didn't change: charges = token usage × unit price, auto-deducted from balance, gift balance spent first, no monthly plan. DeepSeek API has no fixed subscription quota — your account balance is your usage cap, and token usage is visible in real time in the console. Top-up and API key creation both happen on platform.deepseek.com.

Item Rule
Official entry platform.deepseek.com console
Top-up Top up balance in the console, as needed
Billing token usage × unit price, gift balance first
Quota No fixed subscription quota, balance is the usage cap
Rate limit v4-flash 2500, v4-pro 500; HTTP 429 when exceeded
API key Create in console; OpenAI / Anthropic formats

On deals: there's no annual discount or free developer quota — the only real "discount" is cache hits. Repeated system prompts and long chat histories hit the cache and cut input cost by dozens of times; it went up this round, but it's still the biggest cost lever. To connect Claude Code, Cursor or Cline, pick OpenAI- or Anthropic-compatible format and fill in the Base URL and key. Need more concurrency? File a capacity ticket — it costs nothing extra. When balance hits zero, the API stops, so top up as you go instead of hoarding.

How to Control Costs After the Hike

Don't fight peak hours — shift every schedulable and cacheable workload to off-peak. Three concrete levers after this hike:

  1. Run batch jobs off-peak: off-peak is half price, so log analysis, batch completion and offline tasks belong in evenings or non-peak afternoons.
  2. Max out cache hit rate: repeated system prompts, multi-turn chats and RAG contexts are your highest-ROI caching targets; cache-hit input is still dozens of times cheaper than cache-miss.
  3. Use third-party coding plans: DeepSeek has no official subscription, but Volcengine, CtCloud and Baidu Qianfan run DeepSeek models on fixed plans with independent pricing — insulated from this peak/off-peak hike.

Is It Still Worth It?

It depends on your volume and timing — heavy users will probably stay, but you have to crush peak-hour calls. This hike shifts the cost of "working-hours" compute onto pay-as-you-go users. If you call heavily during the day, costs climb fast; if you can move work off-peak and get caching right, the off-peak ¥1.5/¥4.5 input/output is still reasonable.

A few concrete moves:

  1. Measure your peak vs off-peak volume share first — don't fixate on unit price alone.
  2. Route peak-hour work to cache hits and the lighter Flash model where possible.
  3. If budget is tight, switch to a fixed plan like Volcengine to lock a monthly cost.
  4. Keep watching https://api-docs.deepseek.com/quick_start/pricing — peak windows and prices can change.

Sources and verification

Last verified

Deepseek Token API FAQ

What exactly is DeepSeek API's new peak/off-peak pricing after the hike, and how do peak vs off-peak hours get billed?

Starting 2026-08-17 00:00 CST, DeepSeek API switched to peak/off-peak pricing: peak hours are 9:00–12:00 and 14:00–18:00, with off-peak at half the peak rate. At peak, v4-flash input is ¥3 and output ¥9, while v4-pro input is ¥9 and output ¥27; off-peak halves each. Cache-hit input also rose — v4-flash ¥0.10 and v4-pro ¥0.30 at peak.

How big is the DeepSeek API price increase for V4-Flash and V4-Pro, and what's the peak-hour markup?

The markup is significant. At peak, v4-flash cache-miss input tripled (¥1→¥3) and output went 4.5× (¥2→¥9); v4-pro input tripled (¥3→¥9) and output went 4.5× (¥6→¥27). Cache-hit input jumped hardest — v4-pro went from ¥0.025 to ¥0.30 at peak, a 12× increase. Off-peak is milder: 1.5× input and 2.25× output.

After the DeepSeek API price hike, how do I use off-peak scheduling and cache hits to cut V4-Flash and V4-Pro costs?

Two levers: shift what you can off-peak and cache what you can. Off-peak is half price, so move log analysis, batch completion and offline jobs to evenings or non-peak afternoons. Design repeated system prompts, multi-turn chats and RAG contexts for cache hits — cache-hit input is still dozens of times cheaper than cache-miss. Do both and real cost drops meaningfully.

Is DeepSeek API still worth using after the price hike, and how do I switch to a fixed third-party coding plan?

Heavy users will likely stay, but you need to know your peak-hour share. If daytime volume is high and budget is tight, the fastest exit is a fixed plan on [Volcengine](/en/huoshan-fangzhou-token-plan), [CtCloud](/en/tianyi-token-plan) or [Baidu Qianfan](/en/baidu-token-plan) — they run DeepSeek models on plan quotas, insulated from this hike. Switching is cheap: swap the Base URL and API key.

Did DeepSeek API top-up, API key setup or concurrency rate limits change with the peak/off-peak pricing?

No. Top-up is still on platform.deepseek.com, billing stays token usage × unit price with gift balance first, and API keys are created in the console with OpenAI and Anthropic formats. Concurrency is still per-account: v4-flash 2500, v4-pro 500, HTTP 429 when exceeded — capacity tickets are free.
AI Token Plan

Global AI Token Comparison

AI Token Plan is not just a collection of vendor links. It places 41+ domestic and international AI platforms into one comparison framework, covering 395+ model entries plus common plan types such as Token Plans, Coding Plans, IDE Tools and LLM APIs.

When you need to compare AI subscription pricing, coding plan quota, API usage cost or official deal entry points, AI Token Plan brings official prices, plan tiers, usage rules, model capabilities and tool integrations into one place, reducing the need to check multiple vendor sites manually.

Data is continuously organized as vendor pricing pages, plan pages and product documentation change, making the homepage a pricing comparison entry point while detail pages explain whether each vendor plan fits individual developers, team purchasing or long-term API usage.

Prices come from each platform's official site and may change at any time; the official price prevails.

© 2026 AI Token Plan · All rights reserved · First published June 18, 2026 · 64 days running · Sitemap