Minimax Token API Pricing, Top-up, Deals & Official Entry
MiniMax API pricing comparison starts with pay-as-you-go billing after applying for an official API key, not with Token Plan credit pools. MiniMax-M3 is the 1M-context native multimodal flagship, permanently 50% off at ¥2.1/¥8.4 per 1M tokens for ≤512K input; MiniMax-M2.7 and highspeed fit production traffic, with cache read at ¥0.42 and cache write at ¥2.625 helping reduce long-context cost, while service_tier priority bills at 1.5×. This page connects deals, API key application, rate limits and integration setup with Speech-2.8, MiniMax-H3 video, image-01, Music-3.0, Agent API and prepaid speech/video packs so you can judge whether MiniMax API fits complex agents, long-context coding or multimodal content generation.
Below is the complete pricing comparison for Minimax Token API, covering 3 plan tiers, core model capabilities, quotas, and official entry. All data is sourced from the official website to help you decide whether it’s worth it.
Last updated:
Minimax Token API Price Comparison: Core Models
Current API flagship `MiniMax-M3`—native multimodal, 1M context, permanently 50% off at ¥2.1/¥8.4 per 1M tokens for ≤512k input and ¥4.2/¥16.8 for >512k; priority tier (service_tier=priority) ¥3.15/¥12.60; cache read ¥0.42 (≤512k) / ¥0.84 (>512k).
Production workhorse `MiniMax-M2.7`—¥2.1 input, ¥8.4 output per 1M tokens, cache read ¥0.42 and write ¥2.625, for agents, long context, and everyday large-scale calls.
Highspeed `MiniMax-M2.7-highspeed`—¥4.2 input, ¥16.8 output per 1M tokens, same quality as M2.7 with faster responses; cache read/write matches M2.7.
Speech-2.8-HD/Turbo—sync T2A ¥2-3.5 per 10k chars, async long-text same; voice design/cloning ¥9.9 per voice on first synthesis.
Video model `MiniMax-H3` (new-generation multimodal video model)—2K ¥0.80/sec, 768P ¥0.50/sec; up to 5 input images free, ¥0.20/image after; Regeneration 768P→2K ¥0.30/sec; Context-IR ¥5.80 input/¥23.00 output per 1M tokens, for content and multimedia workflows.
Minimax Token API Price Comparison: Plans
MiniMax-M3 is a native multimodal frontier coding model with 1M context—permanently 50% off at ¥2.1 input and ¥8.4 output per 1M tokens for ≤512k input, suited to complex agents and long-horizon code.For everyday high-volume cost-sensitive production, route most traffic to M2.7 and switch to M3 only at complex checkpoints.
It is especially suitable for ongoing sessions, long-context applications, tool-using agents, and production services that need stable long-run scaling.
It is effectively a pay-for-speed version, so whether the upgrade is worth it depends on whether latency truly affects business conversion or user experience.
Minimax Token API Price Comparison: Notes
- M3 ≤512k input is permanently 50% off at ¥2.1/¥8.4 (list ¥4.2/¥16.8); >512k input is also 50% off at ¥4.2/¥16.8 (list ¥8.4/¥33.6) with limited availability. Priority tier (service_tier=priority) bills at 1.5× standard (¥3.15/¥12.60).
- M2.7 / M2.7-highspeed standard pricing: ¥2.1 or ¥4.2 input, ¥8.4 or ¥16.8 output per 1M tokens; cache read ¥0.42, write ¥2.625. M3 priority tier (service_tier=priority) bills at 1.5× standard.
- Speech, video, image, and music have separate list prices; Music-3.0 and lyric generation are currently marked as free promos—see pay-as-you-go docs.
- High-volume speech/video also has prepaid packs: HD speech from ¥630 (2M chars/mo), Turbo from ¥360; video packs from ¥7,000 with a credit pool by model/resolution—a third system separate from Token Plan and pay-as-you-go balance.
Minimax Token API Price Comparison: Tools & Integration
Minimax Token API Pricing, Top-up, Deals, Quota, Usage, Setup & Updates
Page published: