In June and July 2026 working sessions, a global games publisher reported that open-weight models resold through a hyperscaler fell short on pricing, caching and performance. In one partner ecosystem we serve, an AI firm already runs 20B+ tokens a day.‡
One gateway to the world's leading open models.
GLM, Kimi, DeepSeek, MiniMax and Tencent Hunyuan behind one account, one key and one OpenAI- and Anthropic-compatible API, at the model vendors' own list prices with discounts, under a Tencent Cloud contract.
Customer briefing · July 2026 · Singapore region pricing throughout
The contract layer, not another model.
Going direct to the frontier labs means 4 self-serve accounts, 4 payment routes, 4 support paths, and no uptime commitment from any of them. TokenHub collapses that into one Tencent Cloud contract: 8 enterprise models from 5 labs, behind one endpoint and one key, with no gateway margin on the token price.
The gallery carries 27 models. The 8 below are the ones enterprise workloads run; each is reachable through the same endpoint and credential, and appears in the same cost report.
| Model | Provider | Strongest at | Context | Input $/1M | Output $/1M | Cached $/1M |
|---|
The same publisher's China-based teams are barred by company policy from overseas LLM APIs, yet need frontier-grade capability. The ask came from the studios first, GLM 5.2 by name; serving it in region under one Tencent Cloud contract is the route their policy allows.
You pay the lab's price. We carry the contract.
TokenHub charges the same list price you would pay the vendor, to the cent, under one Tencent Cloud contract. What you buy is everything around the token: one counterparty, an enforceable SLA, and a second source you hold open rather than a sole source you bet on.
GLM-5.2 at $1.40 / $4.40 / $0.26 is Z.ai's own published rate. Kimi K3 at $3.00 / $15.00 is Moonshot's. DeepSeek's direct route passes through DeepSeek's exact rate, cache included.*
One contract, one invoice, one procurement review, one support path, the same one as the rest of your Tencent Cloud estate. Not 5 self-serve accounts billed to a corporate card.
99.5% monthly availability, with service-credit vouchers of 10% / 25% / 50% of the month's fee as availability falls, written into the master agreement.*
Your router One more upstream in LiteLLM or your own routing: bulk coding, localisation and batch runs move to open models; the hard tail stays on Claude and GPT.
Second source One relationship holds 5 labs open at once. When the leaderboard moves, switching is a model-name change on terms you already have.
Lab-direct A lab's direct quote starts from the same list number.* If one lab wins one workload, run it direct and keep the rest here.
Committed volume As one indicative anchor, a GLM-5.2 workload in the $15K-40K/month range can qualify for committed terms of up to 35% off list; the number for your volume comes from your account team, in writing.†
Where we are not the answer. If your only criterion is the lowest number per million tokens, resold endpoints exist below our price, some at reduced precision or truncated context, most with no published availability commitment at all. If you need one of those and can carry the operational risk, take it. TokenHub is for the workloads that cannot.
Routing status. Today TokenHub is a clean upstream for the router you already operate. Bundled automatic routing ships inside the Enterprise Token Plan. A general auto-router is on the roadmap for H2 with no committed date.
Same model, 5 providers, and a 2.3× spread in what you actually pay.
Below is one model, Tencent Hy3, as listed by 5 providers on an independent routing marketplace. Headline prices sit within cents of each other; blended price, once cache is counted, does not.
| Provider | Headline input | Output | Cache hit rate | Blended input |
|---|---|---|---|---|
| Tencent Cloud | $0.132 | $0.528 | 91.1% | $0.042 |
| NovitaAI | $0.140 | $0.580 | 88.0% | $0.048 |
| DeepInfra | $0.140 | $0.580 | 69.0% | $0.068 |
| GMICloud | $0.129 | $0.534 | 52.7% | $0.078 |
| AtlasCloud | $0.200 | $0.800 | 69.1% | $0.096 |
GMICloud's headline price is lower than ours; its blended price is 86% higher, because barely half its input hits cache. The rate card is not the bill.
Cache hit rate is mostly a property of your prompts, not of the provider: a stable system prompt, stable tool definitions and long shared context all raise it.
TokenHub publishes a cached rate for every language model. On GLM-5.2 it is $0.26 per 1M cached input tokens, the same cached rate as the model vendor's own endpoint, and the console reports cache hit rate per model in real time. The slider below prices your own reuse; the calculator slide prices a full workload.
Agents, support bots and RAG resend the same prefix on every call. TokenHub bills that repeated prefix at the published cached rate, no code change. Move the slider to your workload's reuse.
Output tokens are never cached and always bill at the full output rate. On GLM-5.2 output is 3.1× the input rate, so generative workloads are output-dominated; the calculator slide accounts for both.
Within 4 index points of the frontier, at a fifth of the output price.
One independently-run index, not a blend of vendor-reported numbers. Open it on your phone while we talk.
| Model | On TokenHub | AA Intelligence Index v4.1 | Context | Output $/1M |
|---|
The index is a weighted composite of 9 evaluations. Compare positions, not ratios: 41 does not mean “two-thirds as capable” as 61.
Kimi K3 scores highest of anything we sell and is also among the slowest. GLM-5.2 runs at roughly 215 output tokens/sec. Pick per workload, not per leaderboard.
No model here beats the frontier leader. The argument is that for most production workloads the remaining gap costs more to close than it is worth.
5 steps from signature to first token.
All of it in the Tencent Cloud console. Nothing to deploy, nothing to install, no SDK to swap.
- 1Account
Tencent Cloud account plus identity verification. Existing cloud customers skip this.
- 2Activate
Activate TokenHub and claim the new-user trial quota from Model Gallery.*
- 3Enable models
Choose models in Model Gallery and switch on pay-as-you-go under Online Inference.
- 4Create a key
Pick a region, create an API key, scope it to models, set its quota and IP allowlist.
- 5Point and go
Change the base URL in the config you already have.
# OpenAI-compatible: any harness, any SDK OPENAI_BASE_URL=https://tokenhub-intl.tencentcloudmaas.com/v1 OPENAI_API_KEY=sk-•••••••• # Anthropic-compatible: Claude Code and Claude-style tools ANTHROPIC_BASE_URL=https://tokenhub-intl.tencentcloudmaas.com ANTHROPIC_AUTH_TOKEN=sk-••••••••
The model name is a request parameter: a 5-lab evaluation fits in an afternoon, on one key and one cost report. No lock-in: exit is a base URL change.
Every token measured, attributed and capped.
The question that stalls most rollouts is not which model. It is who is spending what, and what stops it. That gets answered in the console rather than in a spreadsheet.
Entra ID, Okta or any SAML 2.0 identity provider federates console sign-in through CAM. Admin roles map to the groups you already run.†
Only procurement and platform engineering sign in, with delegated admin scoped by CAM role.
Each team gets its own key, scoped to models, with its own quota and rate ceilings, distributed through your secret manager. The controls are the cards below.
Engineers never open a console. Usage and cost report per key, so team attribution is automatic.
Team-level routing stays in the gateway you already operate (LiteLLM, Kong, Azure API Management); TokenHub is one upstream behind it. Every call rides a governed key, not a personal account, which is the practical answer to shadow AI.
Tokens and spend broken out by service, model and API key over any window, and downloadable.
Requests per minute, time-to-first-token, time-per-output-token and cache hit rate, per model, in real time.
Per-service TPM and QPM ceilings and per-key token quotas. A runaway agent hits a wall, not your budget.
Keys scoped to named models and services, IP allowlisting, and a per-key kill switch.
Threshold alerts routed to SMS, email, phone, WeCom or webhook through Tencent Cloud Observability.
Choose Singapore or Guangzhou. Keys, quotas and monitoring are all region-scoped.*
Where your tokens are processed.
Two serving regions today, and no EU or North America endpoint yet. That fact decides which workloads move first.
Singapore is the international endpoint. Guangzhou serves China-based teams with a separate catalogue and price list. Keys, quotas and monitoring are scoped to the region you pick.
Distance adds roughly 0.2 seconds to time to first token from Western Europe. Generation speed is set by the model, not the route. Measure both on the trial quota; the console reports TTFT per model, live.*
An EU serving region is planned for Q4 2026. North America follows. Regional list prices are published at each launch and may differ from the Singapore list.†
Coding assistants, internal tools and batch pipelines run well from Singapore today. Workloads that process EU player or personal data should wait for the EU endpoints, or go through a capacity-planning session to see whether dedicated capacity can be provisioned in a suitable region. Agreeing that split up front avoids a late legal surprise.
Published list price. No gateway margin.
Pay-as-you-go, USD per million tokens, Singapore region. Sort any column. This is the same table published in the Tencent Cloud documentation.
| Model | Provider | Condition | Input | Output | Cached input | Cache saving |
|---|
Put your own numbers through it.
Indicative guidance only, not an offer. Bands are internal planning ranges, not a published rate card. Any commercial terms require Tencent Cloud account-team review and written approval.
3 ways to buy, so each workload runs on terms that fit it.
A prototype and a production support queue should not sit on the same commercial terms. Mix tiers across workloads under one account.
List price per token, no commitment, every model. The default for evaluation and spiky traffic.
A monthly pool drawn down in real time, with a raised TPM ceiling. Pro covers GLM, Kimi, MiniMax and DeepSeek plus auto-routing; Lite is auto-routing only.
Guaranteed peaks and non-shared clusters are scoped case by case with the account team, where the serving region supports them.
3 things we can do this week.
None of them require a contract, and each one produces a number you can check yourself.
Free tokens on your own account, today. Point one existing harness at TokenHub and run your real prompts against the real models.
Send the token profile: volume, prompt shape, latency target. We return a costed comparison against what you pay today.
Once volume is real, we move you off pay-as-you-go onto the tier that fits it, and quote committed terms in writing.