Global AI Token Plan comparison home

Zhipu Token API Pricing, Top-up, Deals & Official Entry

Zhipu GLM API pricing compares official access, deals and API key application from ¥0.8/¥2, covering GLM-5.3, GLM-5.2, GLM-4.7, GLM-5V-Turbo, CodeGeeX, CogView, CogVideoX, AutoGLM model quotas, usage limits, OpenAI/Anthropic-compatible setup, context cache and 50%-priced Batch API billing.

Token APIGLM-5.3

Zhipu Token API Latest Updates

Analysis

Zhipu GLM-5.3 API Pricing Is Live: Price Comparison vs Kimi Token Plan, Top-Up, Official Entry and Whether It's Worth It

Zhipu published GLM-5.3's pay-as-you-go API pricing on 2026-08-21: ¥8 input / ¥28 output per 1M tokens with cache-hit ¥2, the same price as GLM-5.2 with 1M context. This article runs a price comparison between the GLM-5.3 API and Kimi Token Plan — Kimi's four membership tiers run ¥39/¥79/¥159/¥559 monthly, and since metered API and subscription plans are different shapes, we show how to convert fairly. We also cover top-up, discounts (50% Batch, cache hits), the official entry, API key setup, quota and rate limits, and how to choose between GLM-5.3 metered, GLM Coding Plan, and Kimi Token Plan.

Below is the complete pricing comparison for Zhipu Token API, covering 6 plan tiers, core model capabilities, quotas, and official entry. All data is sourced from the official website to help you decide whether it’s worth it.

Last updated

Zhipu Token API Price Comparison: Core Models

GLM-5.3GLM-5.2GLM-5.1GLM-5GLM-5-TurboGLM-4.7GLM-4.7-FlashXGLM-4.6GLM-4.5-AirGLM-4.5-AirXGLM-4-LongGLM-4-FlashX-250414GLM-4.7-FlashGLM-4-Flash-250414GLM-4-PlusGLM-4-Air-250414GLM-4-AirXGLM-4-AssistantGLM-5V-TurboGLM-4.6VGLM-4.6V-FlashXGLM-4.5VGLM-OCRAutoGLM-PhoneGLM-4.1V-Thinking-FlashXGLM-4.6V-FlashGLM-4.1V-Thinking-FlashGLM-4V-FlashGLM-4V-Plus-0111GLM-ImageCogView-4CogView-3-FlashCogVideoX-3CogVideoX-2Vidu Q1Vidu 2CogVideoX-FlashGLM-TTSGLM-TTS-CloneGLM-ASR-2512GLM-RealtimeGLM-4-VoiceEmbedding-3Embedding-2CharGLM-4EmohaaCodeGeeX-4Rerank
GLM-5.3

Newest flagship text model, on the pricing page since 2026-08-21 marked "New": flat ¥8/¥28 (1M context, 128K output) with cache hit ¥2 and free storage; built for complex software engineering and long-horizon agents, coding experience improved 50% over the previous generation, some cyber-security capability on par with Mythos 5.

GLM-5.2

Previous pricing-page flagship text model with 1M context and open-source SOTA coding—more stable long-horizon execution; supports thinking mode, tools, and MCP.

GLM-5.1

Previous-gen flagship text model with 200K context, tiered at [0,32K) ¥6/¥24 and [32K+) ¥8/¥28, cache hit ¥1.3-¥2; supports thinking mode, tools, and MCP—still first-tier long-horizon capability.

GLM-5-Turbo

Text base optimized for complex long tasks and agents with 200K context and strong continuity—list price slightly below GLM-5.2.

GLM-5V-Turbo

Multimodal coding base accepting image/video/file/text with 200K context—for visual agents and frontend replication.

Zhipu Token API Price Comparison: Plans

GLM-5.3
Newest flagshipRecommended
Input
¥8
Output
¥28
1M context; cache hit ¥2 (free storage)
Usage
GLM-5.3 is the newest flagship for complex software engineering and long-horizon agent tasks, with coding experience improved 50% over the previous generation, some cyber-security capability on par with Mythos 5, and a better balance between quality and token efficiency.
Models
GLM-5.3 flat ¥8/¥28, 1M context (128K max output), cache hit ¥2 with free storage; supports thinking mode, streaming, function calling, structured output—ideal for long-context project-level engineering and complex agents.
Highlights
Cache storage is currently free and cache-hit input is only ¥2 per 1M tokens—design caching for long chats, fixed system prompts, and project-scale context.
Priced identically to GLM-5.2, it is the strongest flagship currently on sale on the Open Platform; suits complex agents, long-horizon coding, project-scale delivery, and production backends needing top reasoning.
Best for
Complex software engineering, long-horizon agents, and project-scale delivery
GLM-5.2
Previous flagship
Input
¥8
Output
¥28
1M context; cache hit ¥2 (free storage)
Usage
GLM-5.2 is the current pricing-page flagship text model with usable 1M context for project-scale engineering, open-source SOTA coding, and more stable long-task execution with reliable spec adherence.
Models
GLM-5.2 flat ¥8/¥28, 1M context (128K max output), cache hit ¥2 with free storage; supports thinking mode, streaming, function calling, structured output—ideal for long-context project-level engineering.
Highlights
Cache storage is currently free and cache-hit input is only ¥2 per 1M tokens—design caching for long chats, fixed system prompts, and project-scale context.
Suited to complex agents, long-horizon coding, project-scale delivery, and production backends needing top reasoning; legacy GLM-5.2 calls in Coding Plan now auto-route to GLM-5.3.
Best for
Long-context engineering, complex agents, and high-quality production backends
GLM-4.7
Mainstream
Input
¥2
Output
¥8
<32K input & <0.2 output ratio · longer context/output tiers cost more · cache hit ¥0.4
Usage
GLM-4.7 upgrades chat, reasoning, coding, and agents with 200K context; pricing tiers by input length and output ratio—¥2/¥8 is a common low-band representative price.
Models
Lower list price than the 5.2/5.1 family—fits high-frequency daily calls, medium-complexity coding, and production traffic optimizing cost.
Highlights
Also supports thinking mode, tools, and cache; pair with GLM-4.5-Air in a route-hard-to-4.7, route-light-to-Air strategy.
A default daily coding model in Coding Plan and a common production default for pay-as-you-go API users.
Best for
Cost-effective production traffic, daily coding, and general agents
GLM-4.5-Air
Value
Input
¥0.8
Output
¥2
Low band <32K input · cache hit ¥0.16 · long input up to ¥1.2/¥8
Usage
GLM-4.5-Air is strong on reasoning, coding, and agents with 128K context—among the lowest paid text entry prices on the platform.
Models
Tiered by input length and output ratio—good for light Q&A, batch preprocessing, classification/extraction, and cost-sensitive high-concurrency endpoints.
Highlights
Available on all Coding Plan tiers; in pay-as-you-go setups it works as a default route or stable low-cost alternative beside Flash models.
For workloads with mostly short input and output, actual bills can stay very low over time.
Best for
Light high-frequency calls, cost-sensitive endpoints, and batch preprocessing
GLM-5V-Turbo
Multimodal coding
Input
¥5
Output
¥22
Image/video/file/text multimodal · ≥32K input ¥7/¥26
Usage
GLM-5V-Turbo is Zhipu’s multimodal agent base combining vision and coding with 200K context—for visual coding, UI replication, and agent workflows.
Models
Accepts image, video, file, and text input with tiered input-length pricing—a top vision choice for complex visual reasoning and frontend code generation.
Highlights
Complements text GLM-5.2/5.1: use 5V when code must follow images/video; pure text engineering is cheaper on 5.2/4.7.
Cache-hit input ¥1.2 per 1M tokens (<32K band)—long multimodal sessions also benefit from caching.
Best for
Visual coding, UI understanding, and multimodal agent teams
GLM-4.7-Flash
Free
Input
¥0
Output
¥0
200K context · platform free model · rate limits apply
Usage
GLM-4.7-Flash is the free tier built on GLM-4.7 with 200K context—good for prototypes, demos, and low-frequency trials.
Models
Input, output, and cache are free but subject to platform concurrency and fair-use policies—not for unlimited production scale.
Highlights
Complements other free models like GLM-4-Flash-250414 and GLM-4.6V-Flash by capability fit.
After validation, migrate to paid tiers like GLM-4.7 or GLM-4.5-Air for stable SLA and higher concurrency.
Best for
Prototyping, learning, and low-frequency trials

Zhipu Token API Price Comparison: Notes

  • GLM-5.3 hit the Open Platform pricing page on 2026-08-21 marked "New" at flat ¥8/¥28 (1M context, 128K output) with cache hit ¥2 and free cache storage; officially positioned for complex software engineering and long-horizon agent tasks, with coding experience improved 50% over the previous generation and some cyber-security capability on par with Mythos 5.
  • `entryPrice` and tier prices here are lowest-band representative quotes. GLM-5.2 flat ¥8/¥28 (1M context, cache hit ¥2). GLM-5 (standalone): <32K ¥4/¥18, ≥32K ¥6/¥22. GLM-4.7 tiers by input length and output ratio. GLM-4.5-Air: <32K low band ¥0.8/¥2. Each request bills at its matching band.
  • Free models such as GLM-4.7-Flash, GLM-4-Flash-250414, GLM-4.6V-Flash, GLM-4.1V-Thinking-Flash, GLM-4V-Flash, CogView-3-Flash, and CogVideoX-Flash remain in the API catalog subject to rate and fair-use rules.
  • GLM-Image ¥0.1/request, CogView-4 ¥0.06/request; CogVideoX-3 ¥1/request. GLM-TTS, GLM-4-Voice, etc. are listed under speech on the pricing page.
  • Team Coding Plan overage bills at 90% of API list price via team keys; standard Open Platform API keys charge account balance in real time.

Zhipu Token API Price Comparison: Tools & Integration

OpenAI-compatible APIAnthropic-compatible APIContext CacheBatch APIFunction CallMCP

Zhipu Token API Pricing, Top-up, Deals, Quota, Usage, Setup & Updates

AnalysisPublished Updated

Zhipu GLM-5.3 API Pricing Is Live: Price Comparison vs Kimi Token Plan, Top-Up, Official Entry and Whether It's Worth It

ZhipuGLM-5.3Kimi Token PlanPrice ComparisonAPI PricingCoding Plan

Zhipu published GLM-5.3's pay-as-you-go API pricing on 2026-08-21: ¥8 input / ¥28 output per 1M tokens with cache-hit ¥2, the same price as GLM-5.2 with 1M context. This article runs a price comparison between the GLM-5.3 API and Kimi Token Plan — Kimi's four membership tiers run ¥39/¥79/¥159/¥559 monthly, and since metered API and subscription plans are different shapes, we show how to convert fairly. We also cover top-up, discounts (50% Batch, cache hits), the official entry, API key setup, quota and rate limits, and how to choose between GLM-5.3 metered, GLM Coding Plan, and Kimi Token Plan.

Page published

Zhipu Token API Pricing, Top-up, Deals & FAQ

AI Token Plan

Global AI Token Comparison

AI Token Plan is not just a collection of vendor links. It places 41+ domestic and international AI platforms into one comparison framework, covering 395+ model entries plus common plan types such as Token Plans, Coding Plans, IDE Tools and LLM APIs.

When you need to compare AI subscription pricing, coding plan quota, API usage cost or official deal entry points, AI Token Plan brings official prices, plan tiers, usage rules, model capabilities and tool integrations into one place, reducing the need to check multiple vendor sites manually.

Data is continuously organized as vendor pricing pages, plan pages and product documentation change, making the homepage a pricing comparison entry point while detail pages explain whether each vendor plan fits individual developers, team purchasing or long-term API usage.

Prices come from each platform's official site and may change at any time; the official price prevails.

© 2026 AI Token Plan · All rights reserved · First published June 18, 2026 · 64 days running · Sitemap