Kimi Token Plan Pricing, Top-up, Deals & Updates | Article

Kimi Token Plan subscription pricing. 4 tiers (Kimi Code, from ¥39/mo/mo). with 1x–10x Agent credits. Kimi K3 and 2 more models. Kimi CLI, Claude Code, Roo Code and 1 more tools integrations. Plan quotas, perks, and official entry—is it worth it?

41
Platforms
395+
Models
140+
IDE Tools
3
Kimi Token Plan models
¥39/mo
Kimi Token Plan entry price
View
Vendor details

Kimi Latest Updates

Kimi K3 Released: World's First Open 3T-Class Model with Frontier Coding and Knowledge Work

Published Updated
ModelK3Kimi K3Model ReleaseOpen SourceMoE
Summary

On July 16, 2026, Moonshot AI released Kimi K3 — a 2.8T-parameter MoE model and the world's first open 3T-class model. Built on Kimi Delta Attention and Attention Residuals, K3 features native vision and a 1M-token context window, delivering frontier performance across coding, knowledge work, and reasoning. Available now on Kimi.com, Kimi Code, and the Kimi API.

Model Overview

Kimi K3 is a 2.8-trillion-parameter open-source MoE model released by Moonshot AI on July 16, 2026 — the world's first to reach the 3T class (source). With native vision and a 1-million-token context window, K3 marks the entry of open-source models into the ultra-large-scale frontier.

In official evaluations, K3 consistently outperforms all other tested models, though it still trails the strongest proprietary models — Claude Fable 5 and GPT 5.6 Sol.

Architectural Innovations

K3's core innovations are Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), paired with Stable LatentMoE across 896 experts for stable large-scale training (source). Overall scaling efficiency improved ~2.5× over K2.

Kimi Delta Attention (KDA)

KDA optimizes attention information flow across sequence length for efficient scaling. A KDA-aware prefix caching implementation has been contributed to the vLLM community, enabling K3 to serve at highly competitive prices at long contexts.

Attention Residuals (AttnRes)

AttnRes selectively retrieves representations across depth rather than accumulating them uniformly, significantly improving training stability and information utilization in very deep models.

Stable LatentMoE

K3 scales MoE to 896 experts (16 active). Quantile Balancing derives expert allocation from router-score quantiles, eliminating heuristic updates and sensitive hyperparameters. Per-Head Muon optimizes attention heads independently.

Coding Capabilities

K3 excels at long-horizon autonomous coding tasks, from kernel optimization and compiler development to chip design (source).

Kernel Optimization

In head-to-head GPU kernel optimization against Claude Fable 5 and GPT 5.6 Sol, K3 completed profiling, rewriting, and benchmarking kernels across NVIDIA Hopper GPUs and alternative GPGPUs within 24 hours, competing strongly with Fable 5 and significantly outperforming Opus 4.8 and GPT 5.6 Sol.

GPU Compiler Development

K3 built MiniTriton from scratch — a compact Triton-like GPU compiler with its own tile-level IR layer, optimization passes, and PTX codegen pipeline. Roofline benchmarks match or beat Triton and torch.compile.

Chip Design

In a single 48-hour autonomous run, K3 designed a chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm²: 1.46M standard cells, 0.277 MB SRAM, INT4 MAC array, sustaining 8,700+ tokens/s decode throughput.

Coding for Research

In a computational astrophysics test, K3 completed in about two hours what typically takes one to two weeks — reproducing I-Love-Q universal relations, cross-validating 20+ papers, evaluating 300+ equations of state, and generating 3,000+ lines of Python code.

Knowledge Work & Agents

K3 shows significant gains in end-to-end knowledge work, supporting interactive research, visual dashboards, and multimodal content creation (source).

  • Interactive Industry Research: Produces drill-down reports through 120+ rounds of recursive self-improvement, extracting data from 2,800+ web searches
  • Widgets & Dashboard: Kimi Work introduces interactive components and personalized dashboards
  • Video Editing: K3's native multimodal architecture handles clip selection, beat synchronization, and multi-round revisions from 56 sources

Availability & Pricing

K3 is live across Kimi.com, Kimi Work, Kimi Code, and the Kimi API — weights were open-sourced on July 27, 2026 (API platform).

Platform How to access
Kimi.com Use directly in browser
Kimi Work Desktop app v3.1.0+ (Windows / Apple Silicon Mac)
Kimi Code Use /model command in terminal
Kimi API Select kimi-k3 model on the platform

API Pricing

Item Price
Cache-hit input $0.30 / MTok
Cache-miss input $3.00 / MTok
Output $15.00 / MTok

With Mooncake's disaggregated inference architecture, cache hit rates exceed 90% on coding workloads.

Limitations

K3's main limitations are sensitivity to thinking history and excessive proactiveness — use explicit behavioral constraints in system prompts (source).

  1. Thinking history sensitivity: K3 was trained in preserved thinking history mode. Agent harnesses must pass back all historical thinking content. Switching models mid-session may cause quality instability.
  2. Excessive proactiveness: K3's training emphasizes long-horizon tasks. It may make unexpected decisions with minor issues or ambiguous intent.
  3. Overall user experience still shows a gap compared to Claude Fable 5 and GPT 5.6 Sol.

Sources and verification

Last verified

Kimi Token Plan FAQ

How much better is Kimi K3 compared to Kimi K2?

K3 achieves approximately 2.5× improvement in overall scaling efficiency over K2. Architecturally, it introduces Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), and scales MoE experts to 896 (16 active), paired with the Stable LatentMoE framework for stable training. K3 substantially outperforms K2 on coding benchmarks like DeepSWE, Terminal-Bench, and SWE Marathon.

What is the API pricing for Kimi K3?

Kimi K3 API pricing: $0.30/MTok for cache-hit input, $3.00/MTok for cache-miss input, and $15.00/MTok for output. Powered by Mooncake's disaggregated inference architecture, the official API achieves a cache hit rate above 90% on coding workloads.

Is Kimi K3 open source? When will the model weights be released?

Yes, K3 is the world's first open 3T-class model. Full model weights were released on July 27, 2026.
AI Token Plan

Global AI Token Comparison

AI Token Plan is not just a collection of vendor links. It places 41+ domestic and international AI platforms into one comparison framework, covering 395+ model entries plus common plan types such as Token Plans, Coding Plans, IDE Tools and LLM APIs.

When you need to compare AI subscription pricing, coding plan quota, API usage cost or official deal entry points, AI Token Plan brings official prices, plan tiers, usage rules, model capabilities and tool integrations into one place, reducing the need to check multiple vendor sites manually.

Data is continuously organized as vendor pricing pages, plan pages and product documentation change, making the homepage a pricing comparison entry point while detail pages explain whether each vendor plan fits individual developers, team purchasing or long-term API usage.

Prices come from each platform's official site and may change at any time; the official price prevails.

© 2026 AI Token Plan · All rights reserved · First published June 18, 2026 · 64 days running · Sitemap