Global AI Token Plan comparison home

Minimax Token API Pricing, Top-up, Deals & Official Entry

MiniMax API pricing comparison starts with pay-as-you-go billing after applying for an official API key, not with Token Plan credit pools. MiniMax-M3 is the 1M-context native multimodal flagship, permanently 50% off at ¥2.1/¥8.4 per 1M tokens for ≤512K input; MiniMax-M2.7 and highspeed fit production traffic, with cache read at ¥0.42 and cache write at ¥2.625 helping reduce long-context cost, while service_tier priority bills at 1.5×. This page connects deals, API key application, rate limits and integration setup with Speech-2.8, MiniMax-H3 video, image-01, Music-3.0, Agent API and prepaid speech/video packs so you can judge whether MiniMax API fits complex agents, long-context coding or multimodal content generation.

Token APIM3 flagship

Minimax Token API Latest Updates

Model

MiniMax H3 Video & Music 3.0 Launch: 2K ¥0.8/sec, ¥1/song — Pricing Comparison and API Key Setup

MiniMax just shipped two new multimodal models: H3 general video (launched 7/31, open-sourced 8/3) and Music 3.0 (launched 8/13). H3 bills per second — 2K ¥0.8/sec, 768P ¥0.5/sec — while Music 3.0 bills per song at ¥1/song with a limited-free tier. This pricing comparison breaks down H3 vs M3/M2.7 text models and Music 3.0 vs Speech-2.8, then covers top-up, deals, official access, quota, usage, rate limits, API key and integration setup to help you decide whether MiniMax is worth it.

Below is the complete pricing comparison for Minimax Token API, covering 3 plan tiers, core model capabilities, quotas, and official entry. All data is sourced from the official website to help you decide whether it’s worth it.

Last updated

Minimax Token API Price Comparison: Core Models

MiniMax-M3MiniMax-M2.7MiniMax-M2.7-highspeedSpeech-2.8MiniMax-H3image-01Music-3.0
MiniMax-M3

Current API flagship `MiniMax-M3`—native multimodal, 1M context, permanently 50% off at ¥2.1/¥8.4 per 1M tokens for ≤512k input and ¥4.2/¥16.8 for >512k; priority tier (service_tier=priority) ¥3.15/¥12.60; cache read ¥0.42 (≤512k) / ¥0.84 (>512k).

MiniMax-M2.7

Production workhorse `MiniMax-M2.7`—¥2.1 input, ¥8.4 output per 1M tokens, cache read ¥0.42 and write ¥2.625, for agents, long context, and everyday large-scale calls.

MiniMax-M2.7-highspeed

Highspeed `MiniMax-M2.7-highspeed`—¥4.2 input, ¥16.8 output per 1M tokens, same quality as M2.7 with faster responses; cache read/write matches M2.7.

Speech-2.8

Speech-2.8-HD/Turbo—sync T2A ¥2-3.5 per 10k chars, async long-text same; voice design/cloning ¥9.9 per voice on first synthesis.

MiniMax-H3

Video model `MiniMax-H3` (new-generation multimodal video model)—2K ¥0.80/sec, 768P ¥0.50/sec; up to 5 input images free, ¥0.20/image after; Regeneration 768P→2K ¥0.30/sec; Context-IR ¥5.80 input/¥23.00 output per 1M tokens, for content and multimedia workflows.

Minimax Token API Price Comparison: Plans

MiniMax-M3
Flagship multimodal
Input
¥2.1
Output
¥8.4
≤512k input · cache read ¥0.42 (>512k ¥0.84) · permanent 50% off · ¥2.1/¥8.4
Usage
MiniMax-M3 is a native multimodal frontier coding model with 1M context—permanently 50% off at ¥2.1 input and ¥8.4 output per 1M tokens for ≤512k input, suited to complex agents and long-horizon code.
Models
Supports ToolCalls, interleaved thinking, and multimodal input; input over 512k bills at 50% off ¥4.2/¥16.8 with limited availability; service_tier=priority accelerates high-concurrency scenarios at 1.5× standard (¥3.15/¥12.60).
Highlights
Cache hits can materially reduce cost; for high-complexity tasks with large per-call consumption, M3 fits critical-path routing better than as a universal default.
For everyday high-volume cost-sensitive production, route most traffic to M2.7 and switch to M3 only at complex checkpoints.
Best for
Teams running complex agents, long-horizon code, and high-value multimodal tasks
MiniMax-M2.7
Latest mainstreamRecommended
Input
¥2.1
Output
¥8.4
Cache read ¥0.42 · write ¥2.625 · production workhorse
Usage
Input, output, and cache read/write are billed separately, which makes it well suited to long-context and repeated-context workloads where caching can materially improve cost control.
Models
The model is MiniMax-M2.7, which works well as a production workhorse for agents, multimodal workflows, and most business tasks that need balanced capability.
Highlights
If you want strong results, long-term cost control, and a unified interface or workflow, this tier is usually more balanced than jumping directly to the highspeed version.
It is especially suitable for ongoing sessions, long-context applications, tool-using agents, and production services that need stable long-run scaling.
Best for
Teams running cost-effective production traffic, long-context applications, and agent workflows
MiniMax-M2.7-highspeed
Highspeed
Input
¥4.2
Output
¥16.8
Higher-speed tier · cache read/write still evaluated separately
Usage
This tier emphasizes speed and throughput rather than materially changing output quality, so it is better for workloads where responsiveness and interactivity rank higher.
Models
MiniMax-M2.7-highspeed is especially suitable for high-concurrency endpoints, interactive products, real-time assistants, and UX flows that are more sensitive to waiting time.
Highlights
If you have already confirmed that M2.7 quality is sufficient but user scale or interaction rhythm turns normal speed into a bottleneck, the highspeed tier becomes the natural upgrade path.
It is effectively a pay-for-speed version, so whether the upgrade is worth it depends on whether latency truly affects business conversion or user experience.
Best for
Teams building high-concurrency interactive products, real-time assistants, and speed-sensitive services

Minimax Token API Price Comparison: Notes

  • M3 ≤512k input is permanently 50% off at ¥2.1/¥8.4 (list ¥4.2/¥16.8); >512k input is also 50% off at ¥4.2/¥16.8 (list ¥8.4/¥33.6) with limited availability. Priority tier (service_tier=priority) bills at 1.5× standard (¥3.15/¥12.60).
  • M2.7 / M2.7-highspeed standard pricing: ¥2.1 or ¥4.2 input, ¥8.4 or ¥16.8 output per 1M tokens; cache read ¥0.42, write ¥2.625. M3 priority tier (service_tier=priority) bills at 1.5× standard.
  • Speech, video, image, and music have separate list prices; Music-3.0 and lyric generation are currently marked as free promos—see pay-as-you-go docs.
  • High-volume speech/video also has prepaid packs: HD speech from ¥630 (2M chars/mo), Turbo from ¥360; video packs from ¥7,000 with a credit pool by model/resolution—a third system separate from Token Plan and pay-as-you-go balance.

Minimax Token API Price Comparison: Tools & Integration

OpenAI-style APICacheAgent

Minimax Token API Pricing, Top-up, Deals, Quota, Usage, Setup & Updates

Page published

Minimax Token API Pricing, Top-up, Deals & FAQ

AI Token Plan

Global AI Token Comparison

AI Token Plan is not just a collection of vendor links. It places 41+ domestic and international AI platforms into one comparison framework, covering 395+ model entries plus common plan types such as Token Plans, Coding Plans, IDE Tools and LLM APIs.

When you need to compare AI subscription pricing, coding plan quota, API usage cost or official deal entry points, AI Token Plan brings official prices, plan tiers, usage rules, model capabilities and tool integrations into one place, reducing the need to check multiple vendor sites manually.

Data is continuously organized as vendor pricing pages, plan pages and product documentation change, making the homepage a pricing comparison entry point while detail pages explain whether each vendor plan fits individual developers, team purchasing or long-term API usage.

Prices come from each platform's official site and may change at any time; the official price prevails.

© 2026 AI Token Plan · All rights reserved · First published June 18, 2026 · 64 days running · Sitemap