LLM Service TokenHub

One gateway to the world's leading open models.

GLM, Kimi, DeepSeek, MiniMax and Tencent Hunyuan behind one account, one key and one OpenAI- and Anthropic-compatible API, at the model vendors' own list prices with discounts, under a Tencent Cloud contract.

27
models in the gallery
8
enterprise models
99.5%
SLA, with service credits

Customer briefing · July 2026 · Singapore region pricing throughout

What it is

The contract layer, not another model.

Going direct to the frontier labs means 4 self-serve accounts, 4 payment routes, 4 support paths, and no uptime commitment from any of them. TokenHub collapses that into one Tencent Cloud contract: 8 enterprise models from 5 labs, behind one endpoint and one key, with no gateway margin on the token price.

1 accountevery model, switchable by one string 2 protocolsOpenAI and Anthropic compatible, zero rewrite Enforceable terms99.5% monthly availability with service credits*

The gallery carries 27 models. The 8 below are the ones enterprise workloads run; each is reachable through the same endpoint and credential, and appears in the same cost report.

ModelProviderStrongest atContext Input $/1MOutput $/1MCached $/1M
Motive one: cost control

In June and July 2026 working sessions, a global games publisher reported that open-weight models resold through a hyperscaler fell short on pricing, caching and performance. In one partner ecosystem we serve, an AI firm already runs 20B+ tokens a day.

Motive two: compliant China access

The same publisher's China-based teams are barred by company policy from overseas LLM APIs, yet need frontier-grade capability. The ask came from the studios first, GLM 5.2 by name; serving it in region under one Tencent Cloud contract is the route their policy allows.

* Availability commitment and compensation schedule: Tencent Cloud SLA, document 301/78868, referenced from the TokenHub SLA (1300/78984). Per-model terms may differ; free-trial and beta usage are excluded. * Tencent Cloud list pricing, Singapore region, USD per million tokens, as published 20 July 2026; full table on the pricing slide. Model Gallery count: Singapore region, 27 July 2026. Guangzhou carries a separate catalogue. Context windows as published by each model provider. ‡ Customer working sessions, June to July 2026; volumes as reported by the companies concerned.
Positioning & fit

You pay the lab's price. We carry the contract.

TokenHub charges the same list price you would pay the vendor, to the cent, under one Tencent Cloud contract. What you buy is everything around the token: one counterparty, an enforceable SLA, and a second source you hold open rather than a sole source you bet on.

No gateway margin
Vendor list price, to the cent

GLM-5.2 at $1.40 / $4.40 / $0.26 is Z.ai's own published rate. Kimi K3 at $3.00 / $15.00 is Moonshot's. DeepSeek's direct route passes through DeepSeek's exact rate, cache included.*

One counterparty
5 labs, one master agreement

One contract, one invoice, one procurement review, one support path, the same one as the rest of your Tencent Cloud estate. Not 5 self-serve accounts billed to a corporate card.

Enforceable
SLA with service credits

99.5% monthly availability, with service-credit vouchers of 10% / 25% / 50% of the month's fee as availability falls, written into the master agreement.*

Your router One more upstream in LiteLLM or your own routing: bulk coding, localisation and batch runs move to open models; the hard tail stays on Claude and GPT.

Second source One relationship holds 5 labs open at once. When the leaderboard moves, switching is a model-name change on terms you already have.

Lab-direct A lab's direct quote starts from the same list number.* If one lab wins one workload, run it direct and keep the rest here.

Committed volume As one indicative anchor, a GLM-5.2 workload in the $15K-40K/month range can qualify for committed terms of up to 35% off list; the number for your volume comes from your account team, in writing.

Where we are not the answer. If your only criterion is the lowest number per million tokens, resold endpoints exist below our price, some at reduced precision or truncated context, most with no published availability commitment at all. If you need one of those and can carry the operational risk, take it. TokenHub is for the workloads that cannot.

Routing status. Today TokenHub is a clean upstream for the router you already operate. Bundled automatic routing ships inside the Enterprise Token Plan. A general auto-router is on the roadmap for H2 with no committed date.

* Vendor list prices compared 27 July 2026 against docs.z.ai, Moonshot's published rates and DeepSeek's official pricing page; parity is asserted per model, not across the whole catalogue. Frontier and open-model rates are the list prices on the capability and pricing slides, snapshots 20 and 27 July 2026; the Elastic tier carries no commitment, and exit is a base-URL change. * Compensation is a service voucher, not cash, capped at the service fee paid for the affected month. Per-model availability terms may differ; free-trial and beta usage are excluded. † Indicative example as of 24 July 2026; not an offer, a rate card or a commitment. Discounts depend on model family, committed volume and term, with account-team review and written approval. Calculator bands are conservative guidance for uncommitted spend; this anchor is a committed-term ceiling.
The number that actually moves your bill

Same model, 5 providers, and a 2.3× spread in what you actually pay.

Below is one model, Tencent Hy3, as listed by 5 providers on an independent routing marketplace. Headline prices sit within cents of each other; blended price, once cache is counted, does not.

ProviderHeadline inputOutput Cache hit rateBlended input
Tencent Cloud$0.132$0.52891.1%$0.042
NovitaAI$0.140$0.58088.0%$0.048
DeepInfra$0.140$0.58069.0%$0.068
GMICloud$0.129$0.53452.7%$0.078
AtlasCloud$0.200$0.80069.1%$0.096

GMICloud's headline price is lower than ours; its blended price is 86% higher, because barely half its input hits cache. The rate card is not the bill.

Why this matters to you

Cache hit rate is mostly a property of your prompts, not of the provider: a stable system prompt, stable tool definitions and long shared context all raise it.

TokenHub publishes a cached rate for every language model. On GLM-5.2 it is $0.26 per 1M cached input tokens, the same cached rate as the model vendor's own endpoint, and the console reports cache hit rate per model in real time. The slider below prices your own reuse; the calculator slide prices a full workload.

Agents, support bots and RAG resend the same prefix on every call. TokenHub bills that repeated prefix at the published cached rate, no code change. Move the slider to your workload's reuse.

0% = every call unique95% = tight agent loop

Output tokens are never cached and always bill at the full output rate. On GLM-5.2 output is 3.1× the input rate, so generative workloads are output-dominated; the calculator slide accounts for both.

List input rate
$1.40
per 1M input tokens · GLM-5.2
Your effective rate
$0.60
57% below list at this reuse
* Provider listings on OpenRouter: Tencent Hy3 captured 27 July 2026; Z.ai GLM-5.2 captured 29 July 2026 (33 listings from 29 providers). Prices are each provider's own marketplace listing, not Tencent Cloud list price. * Blended input = cache hit rate × cached rate + (1 − cache hit rate) × input rate. Cache hit rates are observed across all traffic on an endpoint in a rolling window and will differ from any single workload's. The $0.378 average is traffic-weighted across all GLM-5.2 endpoints at capture. * Effective input rate = list × (1 − cache hit) + cached rate × cache hit. Tencent Cloud list pricing, Singapore, published 20 July 2026.
Capability

Within 4 index points of the frontier, at a fifth of the output price.

One independently-run index, not a blend of vendor-reported numbers. Open it on your phone while we talk.

ModelOn TokenHubAA Intelligence Index v4.1 ContextOutput $/1M
Read it this way

The index is a weighted composite of 9 evaluations. Compare positions, not ratios: 41 does not mean “two-thirds as capable” as 61.

Speed is a separate axis

Kimi K3 scores highest of anything we sell and is also among the slowest. GLM-5.2 runs at roughly 215 output tokens/sec. Pick per workload, not per leaderboard.

What we are not claiming

No model here beats the frontier leader. The argument is that for most production workloads the remaining gap costs more to close than it is worth.

* Artificial Analysis Intelligence Index v4.1, artificialanalysis.ai/models, retrieved 27 July 2026; independently run, 9 evaluations. Scores move as models are added. * Output prices: Tencent Cloud Singapore list (20 July 2026) for TokenHub models; providers' own published list otherwise. Kimi K3 weights are not publicly released.
Onboarding

5 steps from signature to first token.

All of it in the Tencent Cloud console. Nothing to deploy, nothing to install, no SDK to swap.

  1. 1
    Account

    Tencent Cloud account plus identity verification. Existing cloud customers skip this.

  2. 2
    Activate

    Activate TokenHub and claim the new-user trial quota from Model Gallery.*

  3. 3
    Enable models

    Choose models in Model Gallery and switch on pay-as-you-go under Online Inference.

  4. 4
    Create a key

    Pick a region, create an API key, scope it to models, set its quota and IP allowlist.

  5. 5
    Point and go

    Change the base URL in the config you already have.

# OpenAI-compatible: any harness, any SDK
OPENAI_BASE_URL=https://tokenhub-intl.tencentcloudmaas.com/v1
OPENAI_API_KEY=sk-••••••••

# Anthropic-compatible: Claude Code and Claude-style tools
ANTHROPIC_BASE_URL=https://tokenhub-intl.tencentcloudmaas.com
ANTHROPIC_AUTH_TOKEN=sk-••••••••

The model name is a request parameter: a 5-lab evaluation fits in an afternoon, on one key and one cost report. No lock-in: exit is a base URL change.

* Trial quota is subject to the current campaign terms published in the console. Integration guides are published per tool under LLM Service TokenHub › AI Tools Integration.
Control

Every token measured, attributed and capped.

The question that stalls most rollouts is not which model. It is who is spending what, and what stops it. That gets answered in the console rather than in a spreadsheet.

How hundreds of developers run on one contract
Your identity provider

Entra ID, Okta or any SAML 2.0 identity provider federates console sign-in through CAM. Admin roles map to the groups you already run.

Platform team, one console

Only procurement and platform engineering sign in, with delegated admin scoped by CAM role.

One scoped key per team

Each team gets its own key, scoped to models, with its own quota and rate ceilings, distributed through your secret manager. The controls are the cards below.

Developers get a URL and a key

Engineers never open a console. Usage and cost report per key, so team attribution is automatic.

Team-level routing stays in the gateway you already operate (LiteLLM, Kong, Azure API Management); TokenHub is one upstream behind it. Every call rides a governed key, not a personal account, which is the practical answer to shadow AI.

Usage & cost attribution

Tokens and spend broken out by service, model and API key over any window, and downloadable.

Live model monitoring

Requests per minute, time-to-first-token, time-per-output-token and cache hit rate, per model, in real time.

Rate limits & quotas

Per-service TPM and QPM ceilings and per-key token quotas. A runaway agent hits a wall, not your budget.

API key governance

Keys scoped to named models and services, IP allowlisting, and a per-key kill switch.

Alarms & notifications

Threshold alerts routed to SMS, email, phone, WeCom or webhook through Tencent Cloud Observability.

Region scoping

Choose Singapore or Guangzhou. Keys, quotas and monitoring are all region-scoped.*

* Data-residency requirements should be confirmed against the current Tencent Cloud data processing agreement for your contracting entity before they go into a security review. † Console user sign-in federates through CAM identity providers using SAML 2.0 on the International console; OIDC covers role-based programmatic access with temporary keys, not console sign-in. Developers need no console account of any kind.
Data geography

Where your tokens are processed.

Two serving regions today, and no EU or North America endpoint yet. That fact decides which workloads move first.

Serving today

Singapore is the international endpoint. Guangzhou serves China-based teams with a separate catalogue and price list. Keys, quotas and monitoring are scoped to the region you pick.

Latency from Europe

Distance adds roughly 0.2 seconds to time to first token from Western Europe. Generation speed is set by the model, not the route. Measure both on the trial quota; the console reports TTFT per model, live.*

On the roadmap

An EU serving region is planned for Q4 2026. North America follows. Regional list prices are published at each launch and may differ from the Singapore list.

How we would scope it with you

Coding assistants, internal tools and batch pipelines run well from Singapore today. Workloads that process EU player or personal data should wait for the EU endpoints, or go through a capacity-planning session to see whether dedicated capacity can be provisioned in a suitable region. Agreeing that split up front avoids a late legal surprise.

* Estimate from typical public-internet round-trip times, Western Europe to Singapore, July 2026; verify from your own network during the trial. The 99.5% commitment covers availability with service credits; latency is monitored per model in the console, not guaranteed. † Roadmap as communicated by the TokenHub team, July 2026; dates are planning targets, not contractual commitments; confirm current status with your account team.
Pricing

Published list price. No gateway margin.

Pay-as-you-go, USD per million tokens, Singapore region. Sort any column. This is the same table published in the Tencent Cloud documentation.

Model Provider Condition Input Output Cached input Cache saving
* Tencent Cloud list pricing, Singapore region, as published 20 July 2026 (tencentcloud.com/document/product/1300/78937). Guangzhou region pricing differs. Verify current rates against that page. * “Direct route” SKUs are fulfilled on the model vendor's own infrastructure at the vendor's published rate, reached through the same endpoint and key and billed on the same invoice. The 99.5% availability commitment applies at the TokenHub platform and API layer; per-model availability terms are published separately, and Direct route availability depends on the model vendor's service. Note the materially lower cached rate against the Tencent-hosted equivalent.
Estimate

Put your own numbers through it.

 

 

TokenHub estimate
≈$0
 
 
Comparator
≈$0
 
Estimated difference
≈$0
 
 
Evaluate List price

 

Indicative guidance only, not an offer. Bands are internal planning ranges, not a published rate card. Any commercial terms require Tencent Cloud account-team review and written approval.

* Tencent Cloud Singapore list (20 July 2026); comparator = published standard tier (27 July 2026, batch and long-context tiers excluded); identical volumes and cache hit rate on both sides. * Non-binding estimates for discussion; not a quotation, proposal or offer; no obligation on Tencent Cloud. Actual charges depend on metered usage.
Commercial model

3 ways to buy, so each workload runs on terms that fit it.

A prototype and a production support queue should not sit on the same commercial terms. Mix tiers across workloads under one account.

Elastic
Pay-as-you-go

List price per token, no commitment, every model. The default for evaluation and spiky traffic.

Best for · pilots and unpredictable load
Budgeted
Enterprise Token Plan

A monthly pool drawn down in real time, with a raised TPM ceiling. Pro covers GLM, Kimi, MiniMax and DeepSeek plus auto-routing; Lite is auto-routing only.

Best for · teams that need a fixed monthly number
Scoped
Reserved throughput and dedicated capacity

Guaranteed peaks and non-shared clusters are scoped case by case with the account team, where the serving region supports them.

Best for · known peaks and isolated workloads
* Pay-as-you-go and the Enterprise and Personal Token Plans as documented on the Tencent Cloud International site, July 2026. Reserved throughput and dedicated capacity are not yet published there; availability, term and pricing are confirmed per region by the TokenHub team before any commitment. * A Personal Token Plan is also available for individual developers and is managed separately from the Enterprise plan.
Next

3 things we can do this week.

None of them require a contract, and each one produces a number you can check yourself.

01
Claim the trial quota

Free tokens on your own account, today. Point one existing harness at TokenHub and run your real prompts against the real models.

02
Bring us one workload

Send the token profile: volume, prompt shape, latency target. We return a costed comparison against what you pay today.

03
Size the commercial tier

Once volume is real, we move you off pay-as-you-go onto the tier that fits it, and quote committed terms in writing.

TokenHub