Models and multipliers
What each model costs, all on this page
The biggest trust killer in this category is quota you can't verify. So this page lays out every coefficient, capability and regional restriction — along with every multiplier change we've ever made.
How the CU is defined
1 CU = 1,000 output tokens on the reference model
Comparing across models needs one unit, or "equivalent switching" means nothing. Input and output are counted separately because their real costs differ by an order of magnitude.
input tokens × input rate + output tokens × output rate
Both per 1,000 tokens. The frontier tier's output rate is fixed at 1.00 — it's the reference, and every other tier is expressed relative to it.
Input is far cheaper than output
The cost of a long context sits in generation, not in feeding material in. Separate rates keep "paste a big chunk of code" from being strangely expensive.
Every call is itemised
The console records input and output tokens, CU charged and a trace ID for each request, exportable as CSV by date range. A calculator is enough to verify it.
Terms are fixed for subscriptions already sold
Any multiplier change applies to new periods only; increases are announced 30 days ahead. That's a database-level constraint — plan versions are immutable.
Every available model
Models within a tier can back each other up. Name a specific model, or name only the tier and let Prism choose inside it.
前沿
最难的那些任务。贵,但该用的时候省不得- 前沿frontier-a可用
- 上下文
- 1050K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥0.40 · ¥0.237 输出 ¥2.00 · ¥1.19 缓存读 ¥0.04 · ¥0.0237 缓存写 ¥0.50 · ¥0.296 - 工具调用
- 读图
- 结构化输出
- 长上下文
- 前沿frontier-b可用
- 上下文
- 1050K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥0.40 · ¥0.237 输出 ¥2.00 · ¥1.19 缓存读 ¥0.04 · ¥0.0237 缓存写 ¥0.50 · ¥0.296 - 工具调用
- 读图
- 结构化输出
- 长上下文
均衡
日常写代码与长对话的默认选择- 均衡balanced-a可用
- 上下文
- 1000K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥0.16 · ¥0.0948 输出 ¥0.84 · ¥0.498 缓存读 ¥0.016 · ¥0.00948 缓存写 ¥0.20 · ¥0.119 - 工具调用
- 读图
- 结构化输出
- 长上下文
- 均衡balanced-b可用
- 上下文
- 1000K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥0.16 · ¥0.0948 输出 ¥0.84 · ¥0.498 缓存读 ¥0.016 · ¥0.00948 缓存写 ¥0.20 · ¥0.119 - 工具调用
- 读图
- 结构化输出
- 长上下文
- 均衡balanced-c可用
- 上下文
- 1000K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥0.16 · ¥0.0948 输出 ¥0.84 · ¥0.498 缓存读 ¥0.016 · ¥0.00948 缓存写 ¥0.20 · ¥0.119 - 工具调用
- 读图
- 结构化输出
- 长上下文
国产
大陆可用,合规路径最短- 国产domestic-a可用
- 上下文
- 128K
- 大陆
- 可用
每百万 token后付费 · 套餐(pro5)输入 ¥0.06 · ¥0.0356 输出 ¥0.36 · ¥0.213 缓存读 ¥0.006 · ¥0.00356 缓存写 ¥0.075 · ¥0.0445 - 工具调用
- 结构化输出
- 国产domestic-b可用
- 上下文
- 128K
- 大陆
- 可用
每百万 token后付费 · 套餐(pro5)输入 ¥0.06 · ¥0.0356 输出 ¥0.36 · ¥0.213 缓存读 ¥0.006 · ¥0.00356 缓存写 ¥0.075 · ¥0.0445 - 工具调用
- 结构化输出
- 国产domestic-c可用
- 上下文
- 128K
- 大陆
- 可用
每百万 token后付费 · 套餐(pro5)输入 ¥0.06 · ¥0.0356 输出 ¥0.36 · ¥0.213 缓存读 ¥0.006 · ¥0.00356 缓存写 ¥0.075 · ¥0.0445 - 工具调用
- 结构化输出
快速
改个措辞、跑个分类,越快越好- 快速fast-a可用· 目录值
- 上下文
- 0K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥0.02 · ¥0.0119 输出 ¥0.12 · ¥0.0711 缓存读 ¥0.002 · ¥0.00119 缓存写 ¥0.025 · ¥0.0148 - 快速fast-b可用· 目录值
- 上下文
- 0K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥0.02 · ¥0.0119 输出 ¥0.12 · ¥0.0711 缓存读 ¥0.002 · ¥0.00119 缓存写 ¥0.025 · ¥0.0148 - 快速fast-c可用· 目录值
- 上下文
- 0K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥0.02 · ¥0.0119 输出 ¥0.12 · ¥0.0711 缓存读 ¥0.002 · ¥0.00119 缓存写 ¥0.025 · ¥0.0148 - 快速fast-d可用· 目录值
- 上下文
- 0K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥0.02 · ¥0.0119 输出 ¥0.12 · ¥0.0711 缓存读 ¥0.002 · ¥0.00119 缓存写 ¥0.025 · ¥0.0148
可直接点名的模型
把下面这些标识直接填进 model 就会锁定这一条线。我们不会替你换成别的模型—— 它此刻不可用时你拿到的是 503 model_unavailable, 而不是一次悄悄的替换。需要自动切换的话用上面那些能力组的标识。
价目按这条线自己的计费系数算, 可能和它所属能力组的不一样——那正是点名的代价与收益: 放弃自动换线,换来按这条线自己的价计费。两栏仍然是同样两个口径: 后付费零买,和套餐内折算。
- 前沿gpt-6-sol可直接调用
- 上下文
- 922K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥1.50 · ¥0.888 输出 ¥7.49 · ¥4.44 缓存读 ¥0.15 · ¥0.0888 缓存写 ¥1.87 · ¥1.11 - 工具调用
- 结构化输出
- 读图
- 长上下文
- responses
- 均衡gpt-6-luna可直接调用
- 上下文
- 922K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥1.50 · ¥0.888 输出 ¥7.49 · ¥4.44 缓存读 ¥0.15 · ¥0.0888 缓存写 ¥1.87 · ¥1.11 - 工具调用
- 结构化输出
- 长上下文
- responses
- 前沿claude-fable-5可直接调用
- 上下文
- 1000K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥18.95 · ¥11.23 输出 ¥94.74 · ¥56.15 缓存读 ¥1.89 · ¥1.12 缓存写 ¥23.68 · ¥14.04 - 工具调用
- 结构化输出
- 读图
- 长上下文
- 深度思考
- 前沿claude-opus-5可直接调用
- 上下文
- 1000K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥9.47 · ¥5.62 输出 ¥47.37 · ¥28.08 缓存读 ¥0.947 · ¥0.562 缓存写 ¥11.84 · ¥7.02 - 工具调用
- 结构化输出
- 读图
- 长上下文
- 深度思考
- 前沿claude-opus-4-8可直接调用
- 上下文
- 1000K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥9.47 · ¥5.62 输出 ¥47.37 · ¥28.08 缓存读 ¥0.947 · ¥0.562 缓存写 ¥11.84 · ¥7.02 - 工具调用
- 结构化输出
- 读图
- 长上下文
- 深度思考
- 均衡claude-sonnet-5可直接调用
- 上下文
- 1000K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥3.79 · ¥2.25 输出 ¥18.95 · ¥11.23 缓存读 ¥0.379 · ¥0.225 缓存写 ¥4.74 · ¥2.81 - 工具调用
- 结构化输出
- 读图
- 长上下文
- 前沿claude-fable-5-1可直接调用
- 上下文
- 1000K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥67.37 · ¥39.93 输出 ¥336.84 · ¥199.65 缓存读 ¥1.68 · ¥0.998 缓存写 ¥84.21 · ¥49.91 - 工具调用
- 结构化输出
- 读图
- 长上下文
- 深度思考
- 均衡grok-4.5可直接调用
- 上下文
- 500K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥3.40 · ¥2.02 输出 ¥10.20 · ¥6.05 缓存读 ¥0.51 · ¥0.302 缓存写 ¥0.20 · ¥0.119 - 工具调用
- 结构化输出
- 长上下文
- 国产glm-5.3-flash可直接调用
- 上下文
- 128K
- 大陆
- 可用
每百万 token后付费 · 套餐(pro5)输入 ¥0.08 · ¥0.0474 输出 ¥0.28 · ¥0.166 缓存读 ¥0.023 · ¥0.0136 缓存写 ¥0.10 · ¥0.0593 - 工具调用
- 结构化输出
- 深度思考
- 均衡grok-4.6可直接调用
- 上下文
- 500K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥3.40 · ¥2.02 输出 ¥10.20 · ¥6.05 缓存读 ¥0.85 · ¥0.504 缓存写 ¥0.20 · ¥0.119 - 工具调用
- 结构化输出
- 长上下文
- 前沿gpt-5.6可直接调用
- 上下文
- 922K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥3.00 · ¥1.78 输出 ¥14.99 · ¥8.88 缓存读 ¥0.30 · ¥0.178 缓存写 ¥3.75 · ¥2.22 - 工具调用
- 结构化输出
- 读图
- 长上下文
- responses
- 前沿gpt-5.5可直接调用
- 上下文
- 1050K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥3.75 · ¥2.22 输出 ¥22.48 · ¥13.33 缓存读 ¥0.375 · ¥0.222 缓存写 ¥0.50 · ¥0.296 - 工具调用
- 结构化输出
- 读图
- 长上下文
- responses
- 前沿gpt-6-astra可直接调用
- 上下文
- 922K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥7.49 · ¥4.44 输出 ¥37.47 · ¥22.21 缓存读 ¥0.749 · ¥0.444 缓存写 ¥9.37 · ¥5.55 - 工具调用
- 结构化输出
- 读图
- 长上下文
- responses
- 均衡gpt-image-2暂不可用
- 上下文
- 1K
- 大陆
- 不提供
每百万 token后付费 · 套餐(pro5)输入 ¥0.16 · ¥0.0948 输出 ¥0.84 · ¥0.498 缓存读 ¥0.016 · ¥0.00948 缓存写 ¥0.20 · ¥0.119 - image
Every available model
Models within a tier can back each other up. Name a specific model, or name only the tier and let Prism choose inside it.
| Model | Tier | Input | Output | Context | Capabilities | Region | Status |
|---|---|---|---|---|---|---|---|
| frontier-a | Frontier | 0.20 | 1.00 | 1050K | Tool useVisionStructured outputLong context | Outside CN only | Operational |
| frontier-b | Frontier | 0.20 | 1.00 | 1050K | Tool useVisionStructured outputLong context | Outside CN only | Operational |
| balanced-a | Balanced | 0.08 | 0.42 | 1000K | Tool useVisionStructured outputLong context | Outside CN only | Operational |
| balanced-b | Balanced | 0.08 | 0.42 | 1000K | Tool useVisionStructured outputLong context | Outside CN only | Operational |
| balanced-c | Balanced | 0.08 | 0.42 | 1000K | Tool useVisionStructured outputLong context | Outside CN only | Operational |
| domestic-a | Domestic flagship | 0.03 | 0.18 | 128K | Tool useStructured output | Available in CN | Operational |
| domestic-b | Domestic flagship | 0.03 | 0.18 | 128K | Tool useStructured output | Available in CN | Operational |
| domestic-c | Domestic flagship | 0.03 | 0.18 | 128K | Tool useStructured output | Available in CN | Operational |
| fast-a | Fast & light | 0.01 | 0.06 | 0K | — | Outside CN only | Operational |
| fast-b | Fast & light | 0.01 | 0.06 | 0K | — | Outside CN only | Operational |
| fast-c | Fast & light | 0.01 | 0.06 | 0K | — | Outside CN only | Operational |
| fast-d | Fast & light | 0.01 | 0.06 | 0K | — | Outside CN only | Operational |
Status comes from real probes every 15 minutes, not a hand-maintained table. The full uptime history lives on the status page, retained 90 days.
Model identifiers are placeholders. At launch this table is driven live by the catalogue service with real model names and current coefficients.
Which tier for which job
We'd rather you didn't overspend. This table is ordered by which tier is cheapest for a given result — not by which tier is most expensive.
| What you're doing | Suggested | Why |
|---|---|---|
| Long agent chains, cross-file refactors | Frontier | One wrong step compounds. Redoing the work costs far more than the money you'd save. |
| Everyday coding and review | Balanced | The default for most people. On this kind of work the quality gap is essentially imperceptible. |
| Long-form Chinese, research digests | Domestic flagship | Genuinely stronger on Chinese material, and cheaper as a side effect. This is selection, not downgrading. |
| Bulk classification, format conversion | Fast & light | The task has no difficulty in it. Paying frontier rates here is pure waste. |
| Completion, inline suggestions | Fast & light | Latency matters more than quality here. The fast one is the correct one. |
| Not sure | Let Prism decide | auto judges each request on its actual difficulty, which beats picking one tier up front. |
Multiplier change history
Publishing only the current value leaves you guessing whether we've changed it. So here's the whole record: every change, when it was announced, when it took effect.
No changes yet — the service is not live, so this table is empty. It is empty because nothing has happened, not because nothing was recorded; the first change gets appended here with both its announcement and effective dates.
A multiplier **increase** is announced 30 days ahead and applies only to new periods — a subscription you already hold is unaffected within its period. Decreases (cheaper for you) take effect immediately.
Regional availability
Models marked "outside CN only" have not completed generative-AI service filing in mainland China and, per regulation, are not offered to users there.
- 01
The console filters by your location
Unavailable models are greyed out with the reason shown — you won't pick one and then have it fail. Prism's routing excludes them automatically too.
- 02
Please state your location accurately
Circumventing regional restrictions to use a model not offered where you are is the user's own responsibility. This is section 4 of the terms of service.
- 03
Regional rules change
Filing status is not static. When a model's availability changes we notify you in advance and let you switch tiers, or refund the unused portion.
You've seen the rates — now see what saves you money
The same quota, routed by Prism according to task difficulty, typically gets more than twice the work done. Rates are the raw material; routing is where the saving happens.