Tengxun Token API Pricing, Top-up, Deals & Updates | Article

Tengxun Token API pricing. Hy3, Hy3 preview and 11 more models. OpenAI-compatible API, TokenHub, Batch inference and 3 more tools integrations. Includes cache tiers and rate limits. API key setup, usage quotas, and official entry—is it worth it?

41
Platforms
395+
Models
140+
IDE Tools
13
Tengxun Token API models
¥1 / ¥4
Tengxun Token API entry price
View
Vendor details

Hunyuan API Latest Updates

Hunyuan Token API Model Guide: Which Model for Text, Translation, Roleplay, Image, Speech, and Vision

Published Updated
AnalysisHunyuan Token APIHy3 APIHunyuan ModelTencent Cloud APIAPI Model Selection
Summary

Hunyuan Token API's six model families — text, translation, roleplay, image, speech, and vision — each use different billing units, so you can't compare token prices directly. This guide breaks down which model to use for each scenario: when to use cache for text, how to choose between the three translation tiers, why roleplay models are same-price across generations, how to estimate costs with the new per-token image billing, and which speech and vision model to pick.

Hunyuan Token API doesn't apply the same pricing formula to every model — its six families (text, translation, roleplay, image, speech, and vision) use fundamentally different billing approaches. Comparing text token prices against image or speech models is just wrong math from the start.

Prices below are Guangzhou region list prices as of 2026-08-12 (unit: per million tokens). Check the TokenHub console for the latest rates.

The bottom line: text models are Hunyuan's main battlefield — most backend integrations go here. Translation, roleplay, image, speech, and vision are specialized vertical models. Estimate costs per-family; don't compare across families.

Text models: Hunyuan's core, focused on caching

Hunyuan's text models come in two generations: GA and preview. The core difference isn't model capability — it's billing structure: GA uses flat per-token pricing (input ¥1/output ¥4) with no context-length tiers; preview uses tiered pricing that increases with longer contexts (≤16K: ¥1.2/¥4, 16K–32K: ¥1.6/¥6.4, 32K+: ¥2/¥8). (source)

Dimension GA Preview
Billing model Flat ¥1/¥4, no length tiers ≤16K ¥1.2/¥4, 16K–32K ¥1.6/¥6.4, 32K+ ¥2/¥8
Cache hit Input drops to ¥0.25 By tier: ¥0.4 / ¥0.6 / ¥0.8
Best for Agents, coding, long context Legacy maintenance

The decision is simple: go GA for new integrations — predictable flat pricing. If you're on preview and your contexts always stay in the shortest tier, migration isn't urgent, but preview will be deprecated.

What really impacts cost isn't GA vs. preview — it's cache hit rate. Hunyuan text models all support automatic context caching. On GA, cache hits drop input from ¥1 to ¥0.25 (75% off); on preview, from ¥1.2 to ¥0.4 (67% off). In scenarios like agents and coding that frequently reuse system prompts and conversation history, caching is your biggest cost lever.

Translation models: Pro and Plus at the same price, Lite lower

Pro and Plus share the same unit price (input ¥0.5/output ¥2); Lite is lower (input ¥0.3/output ¥1.2). These aren't linear quality-vs-price tiers but precision options for different language coverage and terminology needs. (source)

Tier Input / Output Recommended For
Pro ¥0.5 / ¥2 Professional translation pipelines, terminology-intensive docs
Plus ¥0.5 / ¥2 (same as Pro) Multi-language combinations, moderate precision
Lite ¥0.3 / ¥1.2 Daily sentence-level translation, short-text batches

If translation quality is core to your product (localization tools, cross-border documents), evaluate from Pro or Plus. If you're just polishing short texts in bulk, Lite is the best value entry. All three are postpaid per input/output token, using the same account balance as text models.

Roleplay models: same price across generations, go Latest

Hunyuan has two roleplay models: Hy-Role and Hy-Role-Latest. Same price across generations — input ¥2.4/output ¥9.6. No decision fatigue: start with Hy-Role-Latest. These models suit chatbots, virtual characters, and interactive storytelling where consistent persona and tone matter more than general text performance.

Roleplay models have 2.4x higher token unit prices than text models (¥1/¥4) — so don't compare them directly. Agents go to text, chatbots go to roleplay. Different problems, different models.

Image models: switched to per-token billing, cost estimation changed

Hunyuan image models have switched from per-image to per-token billing — ¥10/MT, with ~20K tokens consumed per generated image (at 1024 resolution). You can't estimate with the old "per image" mental model anymore. Token consumption per generated image depends on resolution and complexity. You need to measure actual token usage in your scenario to estimate costs accurately. (source)

Legacy per-image image models and video/3D generation models have moved to the deprecated section. New visual generation integrations should use the new image model or partner models on TokenHub. Image model token pricing is separate from text models — you can't derive image costs from text unit prices.

Speech and vision: independent billing frameworks

Speech recognition (Hy-ASR-3.0-Preview) bills at ¥10/MT, approximately ~22 tokens per second of audio. It's designed for speech-to-text, meeting transcription, and subtitle generation.

Vision understanding models use input/output token billing at higher rates than text models — HY-Vision-2.0-Instruct at ¥7.5/¥17.5, HY-Vision-1.5-Thinking at ¥3/¥9, HY-Vision-Video at ¥3/¥9, YT-VITA at ¥1.2/¥3.5. Actual costs depend on the resolution and quantity of input images. Suited for vision-language understanding and VQA, with routing across models by task complexity.

Both speech and vision token prices aren't directly comparable to text models — they solve different problems with completely different token consumption patterns. Estimate costs independently in your actual usage scenarios.

TL;DR: three questions before picking a model

Question Direction
What's your task type? Text agent → GA; Translation → Pro/Lite; Chatbot → Roleplay; Vision-language → Vision; Speech → Speech recognition
Is your cache hit rate high? High → caching is the #1 cost lever; Low → optimize prompt structure for reuse
What's your deployment region? Guangzhou vs Singapore unit prices differ — compare before cross-region deployment

Hunyuan Token API has one of the most diverse billing ecosystems — six families, six rulesets. Classify your scenario first, then compare within the family. Never compare token unit prices across families.

Sources and verification

Last verified

Tengxun Token API FAQ

How do the six Hunyuan Token API model families differ in billing, and how should I compare text, translation, roleplay, image, speech, and vision pricing?

Hunyuan Token API families use different billing models — text and translation are postpaid per input/output token with cache discounts, roleplay also uses token billing at higher unit rates, image has switched to per-token billing (different from legacy per-image), speech recognition bills separately per token converted from audio duration, and vision understanding uses input/output token billing at the highest rates. Don't compare absolute token prices across families — factor in cache hit rate, Batch mode, and deployment region. Confirm specific pricing on this page and the TokenHub console.

How do I choose between Hunyuan Token API text models? What's the difference between the GA and preview versions?

Text models are the core Hunyuan product line. The GA version uses flat unit pricing with no context-length tiers, while the preview version has tiered pricing that increases with longer contexts — and the preview will eventually be deprecated. The GA version has lower input pricing, making it the preferred choice for agents, coding, and long-context scenarios. Cache hits can significantly reduce input costs, with the biggest gains in scenarios that frequently reuse context.

How do the three Hunyuan Token API translation model tiers — Pro, Plus, and Lite — differ, and how should I choose?

Translation models come in three tiers — Pro and Plus share the same unit price, while Lite is lower. These aren't linear quality vs. price tiers but precision options for different language coverage and terminology scenarios. If translation quality is core to your product, evaluate from Pro or Plus. Lite is the most cost-effective entry for daily sentence-level translation and short-text batches. All three are postpaid per input/output token within the same balance system as text models.

What's the difference between Hunyuan Token API roleplay models Hy-Role and Hy-Role-Latest? Same price, which should I use?

Hy-Role and Hy-Role-Latest are Hunyuan's character roleplay models with identical unit pricing — no price difference between generations. Hy-Role-Latest is the newer generation suited for virtual characters, chatbots, and interactive storytelling requiring consistent persona and tone. New integrations should go directly to Hy-Role-Latest. Model performance should be validated against actual scenario evaluations.

How do I estimate Hunyuan Token API image costs after the switch to per-token billing? What about speech and vision models?

Image models have switched from per-image to per-token billing — cost estimation now requires converting token consumption per generated image rather than the old per-image thinking. Speech recognition bills separately per token, approximately converting from audio seconds to token count. Vision understanding uses input/output token billing at higher rates than text models, with actual costs depending on input image resolution and quantity. None of these three non-text families are directly comparable to language model token pricing. Confirm specific rates and conversion methods on this page and the official pricing docs.
AI Token Plan

Global AI Token Comparison

AI Token Plan is not just a collection of vendor links. It places 41+ domestic and international AI platforms into one comparison framework, covering 395+ model entries plus common plan types such as Token Plans, Coding Plans, IDE Tools and LLM APIs.

When you need to compare AI subscription pricing, coding plan quota, API usage cost or official deal entry points, AI Token Plan brings official prices, plan tiers, usage rules, model capabilities and tool integrations into one place, reducing the need to check multiple vendor sites manually.

Data is continuously organized as vendor pricing pages, plan pages and product documentation change, making the homepage a pricing comparison entry point while detail pages explain whether each vendor plan fits individual developers, team purchasing or long-term API usage.

Prices come from each platform's official site and may change at any time; the official price prevails.

© 2026 AI Token Plan · All rights reserved · First published June 18, 2026 · 64 days running · Sitemap