Global AI Token Plan comparison home

Kimi Token API Pricing, Top-up, Deals & Official Entry

Kimi API pricing comparison should start from Kimi K3 flagship, Kimi K2.7 Code programming, and tiered cache billing: Kimi K3 is ¥20/1M tokens input (¥2 cache hit) and ¥100 output with 1M context and always-on inference; Kimi K2.7 Code standard is ¥1.30 / ¥6.50 / ¥27 cache hit / uncached input / output, HighSpeed at ¥2.60 / ¥13 / ¥54; Kimi K2.6 multimodal is ¥1.10 / ¥6.50 / ¥27. This page connects the Moonshot platform, API key setup, automatic context caching, ToolCalls, JSON Mode, `$web_search` at ¥0.03/call, free file extraction, Claude Code / Cline / Roo Code integration, and Tier0–Tier5 rate limits, with deprecation notes that Kimi K2.5 and Moonshot V1 will go offline by 2026-08-31.

Token APIKimi K3 flagship

Kimi Token API Latest Updates

Product

Kimi Launches Agent Membership: Tiered Pricing, K3 Open-Source Impact, and Hong Kong IPO

In August 2026, Moonshot AI launched the Kimi Agent membership system with four tiers: Andante ($7/mo), Moderato ($14/mo), Allegretto ($28/mo), and Allegro ($97/mo), plus Vivace at $199/mo for international users. Following K3's open-source release that triggered a $3.3 trillion chip stock sell-off, Kimi is accelerating its commercialization with ARR surpassing $300 million and a Hong Kong IPO in preparation.

Below is the complete pricing comparison for Kimi Token API, covering 5 plan tiers, core model capabilities, quotas, and official entry. All data is sourced from the official website to help you decide whether it’s worth it.

Last updated

Kimi Token API Price Comparison: Core Models

Kimi K3Kimi K2.7 CodeKimi K2.7 Code HighSpeedKimi K2.6Kimi K2.5Moonshot V1 128KMoonshot V1 32KMoonshot V1 8KMoonshot V1 128K VisionMoonshot V1 32K VisionMoonshot V1 8K Vision
Kimi K3

Kimi's strongest flagship `kimi-k3`: 2.8T parameters, native vision, 1M-token context, always-on inference, ToolCalls (with `tool_choice` and dynamic tool loading), structured output, context caching. ¥2.00 cache, ¥20.00 uncached input, ¥100.00 output per 1M tokens.

Kimi K2.7 Code / HighSpeed

Current coding workhorse `kimi-k2.7-code` and `kimi-k2.7-code-highspeed`: text, image, and video input in reasoning-only mode with 256k context. Standard ¥1.30 / ¥6.50 / ¥27.00; HighSpeed ~180 tokens/s at ¥2.60 / ¥13.00 / ¥54.00. Built for complex engineering, multi-step agents, and long-horizon refactors.

Kimi K2.6

General multimodal model `kimi-k2.6`: vision/text/video input, reasoning/non-reasoning modes, chat and agent tasks, 256k context. ¥1.10 cache, ¥6.50 uncached input, ¥27.00 output per 1M tokens—suited to comprehensive workflows needing vision and general agent capabilities.

Kimi K2.5

`kimi-k2.5` stopped for new users, going offline by 2026-08-31. Current pricing ¥0.70 / ¥4.00 / ¥21.00 per 1M tokens, 256k context. Existing users should migrate to Kimi K3, Kimi K2.7 Code, or Kimi K2.6 immediately.

Moonshot V1

Moonshot V1 series stopped for new users, going offline by 2026-08-31. Includes 8K (¥2/¥10), 32K (¥5/¥20), 128K (¥10/¥30) text and Vision Preview. Migration: short text/vision → Kimi K2.6, long text → Kimi K3 (1M context), lightweight high-frequency → Kimi K2.7 Code standard.

Kimi Token API Price Comparison: Plans

Kimi K3
Ultimate flagshipRecommended
Input
¥20
Output
¥100
¥2.00 cache hit · ¥20.00 uncached input · ¥100.00 output
Usage
kimi-k3 is Kimi's most capable model to date—2.8 trillion parameters, native visual understanding, 1M-token context window, always-on inference, and adjustable reasoning_effort.
Models
Built for long-horizon coding and end-to-end knowledge work, with ToolCalls (including tool_choice constraints and dynamic tool loading), JSON Mode, Structured Output (response_format / JSON Schema), Partial Mode, automatic context caching, and web search.
Highlights
Cached input at ¥2.00/1M tokens (uncached ¥20.00), output ¥100.00; 1M context and always-on inference make it ideal for high-value, complex, low-tolerance core tasks.
Best for
Complex software engineering, knowledge-intensive agents, ultra-long-context reasoning, and critical tasks requiring peak model capability
Kimi K2.7 Code
Coding flagship
Input
¥6.5
Output
¥27
Standard ¥1.30 / ¥6.50 / ¥27 cache hit / input / output; HighSpeed ¥2.60 / ¥13 / ¥54
Usage
kimi-k2.7-code is Kimi's current coding model, supporting text, image, and video input in reasoning-only mode with 256k context; more reliable instruction following in long contexts at higher coding success rates.
Models
kimi-k2.7-code-highspeed shares the same model with ~180 tokens/s output (up to 260 tokens/s in short context) at 2x pricing: ¥2.60 cache, ¥13 uncached input, ¥54 output per 1M tokens.
Highlights
Supports automatic context caching, ToolCalls, JSON Mode, Partial Mode, etc.; lower input cost than Kimi K3 makes it the everyday workhorse for coding agents.
Best for
AI coding tools, complex code agents, long-context engineering, and dev teams needing faster output
Kimi K2.6
Multimodal general
Input
¥6.5
Output
¥27
¥1.10 cache hit · ¥6.50 uncached input · ¥27.00 output
Usage
kimi-k2.6 is a general multimodal model with text, image, and video input, reasoning and non-reasoning modes, chat and agent tasks, and 256k context; stable instruction following and self-correction.
Models
Covers ToolCalls, JSON Mode, Partial Mode, automatic context caching, and web search; input price equals Kimi K2.7 Code standard but optimised for general tasks rather than programming specifically.
Highlights
If your workflow needs both vision input and general agent tasks, Kimi K2.6 is a more versatile entry point than Kimi K2.7 Code; pure coding and long-context engineering still prefer Kimi K2.7 Code or Kimi K3.
Best for
Comprehensive scenarios needing visual understanding, general agent tasks, and multimodal input
Kimi K2.5
Deprecating
Input
¥4
Output
¥21
¥0.70 cache hit · ¥4.00 uncached input · ¥21.00 output · Stopped for new users, offline Aug 31
Usage
kimi-k2.5 has stopped accepting new users and will go offline across the platform by 2026-08-31. Teams still using Kimi K2.5 should migrate to Kimi K3 or Kimi K2.7 Code as soon as possible.
Models
Migration advice: for general agent and long-text, prefer Kimi K2.6 (similar price at ¥6.50 / ¥27, cache ¥1.10 vs Kimi K2.5's ¥0.70); for coding use Kimi K2.7 Code; for critical tasks use Kimi K3.
Best for
Existing teams still using this model—migrate now
Moonshot V1
End-of-life
Input
¥2
Output
¥10
8K ¥2/¥10 · 32K ¥5/¥20 · 128K ¥10/¥30 · Offline Aug 31
Usage
Moonshot V1 classic generation models (8K/32K/128K and corresponding Vision Preview) have stopped accepting new users and will go offline by 2026-08-31.
Models
Migration advice: for short text and image understanding, switch to Kimi K2.6 (more comprehensive vision); for long text, prefer Kimi K3 (1M context far exceeds V1 128K); for lightweight high-frequency, evaluate Kimi K2.7 Code standard with cache-hit optimisation.
Best for
Existing users must migrate to Kimi K3 / Kimi K2.7 Code / Kimi K2.6 immediately

Kimi Token API Price Comparison: Notes

  • Kimi K3 flagship: ¥2.00 cache hit, ¥20.00 uncached input, ¥100.00 output per 1M tokens, 1,048,576-token context, always-on inference with configurable `reasoning_effort` (low / high / max, default max).
  • Kimi K2.7 Code standard: ¥1.30 / ¥6.50 / ¥27.00 (cache/input/output). HighSpeed same model faster (~180 tokens/s): ¥2.60 / ¥13.00 / ¥54.00. Kimi K2.6 multimodal: ¥1.10 / ¥6.50 / ¥27.00.
  • Kimi K2.5 (¥0.70 / ¥4.00 / ¥21.00) and Moonshot V1 have stopped accepting new users and will go offline by 2026-08-31. Migrate to Kimi K3 or Kimi K2.7 Code promptly.
  • Each successful `$web_search` trigger costs an extra ¥0.03; file extraction and storage APIs are temporarily free, but extracted document content billed as model input tokens.

Kimi Token API Price Comparison: Tools & Integration

OpenAI-compatible APIClaude CodeClineRoo Code

Kimi Token API Pricing, Top-up, Deals, Quota, Usage, Setup & Updates

Page published

Kimi Token API Pricing, Top-up, Deals & FAQ

AI Token Plan

Global AI Token Comparison

AI Token Plan is not just a collection of vendor links. It places 41+ domestic and international AI platforms into one comparison framework, covering 395+ model entries plus common plan types such as Token Plans, Coding Plans, IDE Tools and LLM APIs.

When you need to compare AI subscription pricing, coding plan quota, API usage cost or official deal entry points, AI Token Plan brings official prices, plan tiers, usage rules, model capabilities and tool integrations into one place, reducing the need to check multiple vendor sites manually.

Data is continuously organized as vendor pricing pages, plan pages and product documentation change, making the homepage a pricing comparison entry point while detail pages explain whether each vendor plan fits individual developers, team purchasing or long-term API usage.

Prices come from each platform's official site and may change at any time; the official price prevails.

© 2026 AI Token Plan · All rights reserved · First published June 18, 2026 · 64 days running · Sitemap