Zhipu Token API Pricing, Top-up, Deals & Official Entry
Zhipu GLM API pricing compares official access, deals and API key application from ¥0.8/¥2, covering GLM-5.3, GLM-5.2, GLM-4.7, GLM-5V-Turbo, CodeGeeX, CogView, CogVideoX, AutoGLM model quotas, usage limits, OpenAI/Anthropic-compatible setup, context cache and 50%-priced Batch API billing.
Below is the complete pricing comparison for Zhipu Token API, covering 6 plan tiers, core model capabilities, quotas, and official entry. All data is sourced from the official website to help you decide whether it’s worth it.
Last updated:
Zhipu Token API Price Comparison: Core Models
Newest flagship text model, on the pricing page since 2026-08-21 marked "New": flat ¥8/¥28 (1M context, 128K output) with cache hit ¥2 and free storage; built for complex software engineering and long-horizon agents, coding experience improved 50% over the previous generation, some cyber-security capability on par with Mythos 5.
Previous pricing-page flagship text model with 1M context and open-source SOTA coding—more stable long-horizon execution; supports thinking mode, tools, and MCP.
Previous-gen flagship text model with 200K context, tiered at [0,32K) ¥6/¥24 and [32K+) ¥8/¥28, cache hit ¥1.3-¥2; supports thinking mode, tools, and MCP—still first-tier long-horizon capability.
Text base optimized for complex long tasks and agents with 200K context and strong continuity—list price slightly below GLM-5.2.
Multimodal coding base accepting image/video/file/text with 200K context—for visual agents and frontend replication.
Zhipu Token API Price Comparison: Plans
Priced identically to GLM-5.2, it is the strongest flagship currently on sale on the Open Platform; suits complex agents, long-horizon coding, project-scale delivery, and production backends needing top reasoning.
Suited to complex agents, long-horizon coding, project-scale delivery, and production backends needing top reasoning; legacy GLM-5.2 calls in Coding Plan now auto-route to GLM-5.3.
A default daily coding model in Coding Plan and a common production default for pay-as-you-go API users.
For workloads with mostly short input and output, actual bills can stay very low over time.
Cache-hit input ¥1.2 per 1M tokens (<32K band)—long multimodal sessions also benefit from caching.
After validation, migrate to paid tiers like GLM-4.7 or GLM-4.5-Air for stable SLA and higher concurrency.
Zhipu Token API Price Comparison: Notes
- GLM-5.3 hit the Open Platform pricing page on 2026-08-21 marked "New" at flat ¥8/¥28 (1M context, 128K output) with cache hit ¥2 and free cache storage; officially positioned for complex software engineering and long-horizon agent tasks, with coding experience improved 50% over the previous generation and some cyber-security capability on par with Mythos 5.
- `entryPrice` and tier prices here are lowest-band representative quotes. GLM-5.2 flat ¥8/¥28 (1M context, cache hit ¥2). GLM-5 (standalone): <32K ¥4/¥18, ≥32K ¥6/¥22. GLM-4.7 tiers by input length and output ratio. GLM-4.5-Air: <32K low band ¥0.8/¥2. Each request bills at its matching band.
- Free models such as GLM-4.7-Flash, GLM-4-Flash-250414, GLM-4.6V-Flash, GLM-4.1V-Thinking-Flash, GLM-4V-Flash, CogView-3-Flash, and CogVideoX-Flash remain in the API catalog subject to rate and fair-use rules.
- GLM-Image ¥0.1/request, CogView-4 ¥0.06/request; CogVideoX-3 ¥1/request. GLM-TTS, GLM-4-Voice, etc. are listed under speech on the pricing page.
- Team Coding Plan overage bills at 90% of API list price via team keys; standard Open Platform API keys charge account balance in real time.
Zhipu Token API Price Comparison: Tools & Integration
Zhipu Token API Pricing, Top-up, Deals, Quota, Usage, Setup & Updates
Page published: