Models and multipliers

What each model costs, all on this page

The biggest trust killer in this category is quota you can't verify. So this page lays out every coefficient, capability and regional restriction — along with every multiplier change we've ever made.

How the CU is defined

1 CU = 1,000 output tokens on the reference model

Comparing across models needs one unit, or "equivalent switching" means nothing. Input and output are counted separately because their real costs differ by an order of magnitude.

The formula

input tokens × input rate + output tokens × output rate

Both per 1,000 tokens. The frontier tier's output rate is fixed at 1.00 — it's the reference, and every other tier is expressed relative to it.

Why separate

Input is far cheaper than output

The cost of a long context sits in generation, not in feeding material in. Separate rates keep "paste a big chunk of code" from being strangely expensive.

How to check

Every call is itemised

The console records input and output tokens, CU charged and a trace ID for each request, exportable as CSV by date range. A calculator is enough to verify it.

What won't move

Terms are fixed for subscriptions already sold

Any multiplier change applies to new periods only; increases are announced 30 days ahead. That's a database-level constraint — plan versions are immutable.

Every available model

Models within a tier can back each other up. Name a specific model, or name only the tier and let Prism choose inside it.

前沿

最难的那些任务。贵,但该用的时候省不得
2/2 可用
  • frontier-a
    可用
    前沿
    上下文
    1050K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥0.40·¥0.237
    输出¥2.00·¥1.19
    缓存读¥0.04·¥0.0237
    缓存写¥0.50·¥0.296
    • 工具调用
    • 读图
    • 结构化输出
    • 长上下文
  • frontier-b
    可用
    前沿
    上下文
    1050K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥0.40·¥0.237
    输出¥2.00·¥1.19
    缓存读¥0.04·¥0.0237
    缓存写¥0.50·¥0.296
    • 工具调用
    • 读图
    • 结构化输出
    • 长上下文

均衡

日常写代码与长对话的默认选择
3/3 可用
  • balanced-a
    可用
    均衡
    上下文
    1000K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥0.16·¥0.0948
    输出¥0.84·¥0.498
    缓存读¥0.016·¥0.00948
    缓存写¥0.20·¥0.119
    • 工具调用
    • 读图
    • 结构化输出
    • 长上下文
  • balanced-b
    可用
    均衡
    上下文
    1000K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥0.16·¥0.0948
    输出¥0.84·¥0.498
    缓存读¥0.016·¥0.00948
    缓存写¥0.20·¥0.119
    • 工具调用
    • 读图
    • 结构化输出
    • 长上下文
  • balanced-c
    可用
    均衡
    上下文
    1000K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥0.16·¥0.0948
    输出¥0.84·¥0.498
    缓存读¥0.016·¥0.00948
    缓存写¥0.20·¥0.119
    • 工具调用
    • 读图
    • 结构化输出
    • 长上下文

国产

大陆可用,合规路径最短
3/3 可用
  • domestic-a
    可用
    国产
    上下文
    128K
    大陆
    可用
    每百万 token后付费 · 套餐(pro5)
    输入¥0.06·¥0.0356
    输出¥0.36·¥0.213
    缓存读¥0.006·¥0.00356
    缓存写¥0.075·¥0.0445
    • 工具调用
    • 结构化输出
  • domestic-b
    可用
    国产
    上下文
    128K
    大陆
    可用
    每百万 token后付费 · 套餐(pro5)
    输入¥0.06·¥0.0356
    输出¥0.36·¥0.213
    缓存读¥0.006·¥0.00356
    缓存写¥0.075·¥0.0445
    • 工具调用
    • 结构化输出
  • domestic-c
    可用
    国产
    上下文
    128K
    大陆
    可用
    每百万 token后付费 · 套餐(pro5)
    输入¥0.06·¥0.0356
    输出¥0.36·¥0.213
    缓存读¥0.006·¥0.00356
    缓存写¥0.075·¥0.0445
    • 工具调用
    • 结构化输出

快速

改个措辞、跑个分类,越快越好
4/4 可用
  • fast-a
    可用· 目录值
    快速
    上下文
    0K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥0.02·¥0.0119
    输出¥0.12·¥0.0711
    缓存读¥0.002·¥0.00119
    缓存写¥0.025·¥0.0148
  • fast-b
    可用· 目录值
    快速
    上下文
    0K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥0.02·¥0.0119
    输出¥0.12·¥0.0711
    缓存读¥0.002·¥0.00119
    缓存写¥0.025·¥0.0148
  • fast-c
    可用· 目录值
    快速
    上下文
    0K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥0.02·¥0.0119
    输出¥0.12·¥0.0711
    缓存读¥0.002·¥0.00119
    缓存写¥0.025·¥0.0148
  • fast-d
    可用· 目录值
    快速
    上下文
    0K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥0.02·¥0.0119
    输出¥0.12·¥0.0711
    缓存读¥0.002·¥0.00119
    缓存写¥0.025·¥0.0148

可直接点名的模型

把下面这些标识直接填进 model 就会锁定这一条线。我们不会替你换成别的模型—— 它此刻不可用时你拿到的是 503 model_unavailable, 而不是一次悄悄的替换。需要自动切换的话用上面那些能力组的标识。

价目按这条线自己的计费系数算, 可能和它所属能力组的不一样——那正是点名的代价与收益: 放弃自动换线,换来按这条线自己的价计费。两栏仍然是同样两个口径: 后付费零买,和套餐内折算。

  • gpt-6-sol
    可直接调用
    前沿
    上下文
    922K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥1.50·¥0.888
    输出¥7.49·¥4.44
    缓存读¥0.15·¥0.0888
    缓存写¥1.87·¥1.11
    • 工具调用
    • 结构化输出
    • 读图
    • 长上下文
    • responses
  • gpt-6-luna
    可直接调用
    均衡
    上下文
    922K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥1.50·¥0.888
    输出¥7.49·¥4.44
    缓存读¥0.15·¥0.0888
    缓存写¥1.87·¥1.11
    • 工具调用
    • 结构化输出
    • 长上下文
    • responses
  • claude-fable-5
    可直接调用
    前沿
    上下文
    1000K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥18.95·¥11.23
    输出¥94.74·¥56.15
    缓存读¥1.89·¥1.12
    缓存写¥23.68·¥14.04
    • 工具调用
    • 结构化输出
    • 读图
    • 长上下文
    • 深度思考
  • claude-opus-5
    可直接调用
    前沿
    上下文
    1000K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥9.47·¥5.62
    输出¥47.37·¥28.08
    缓存读¥0.947·¥0.562
    缓存写¥11.84·¥7.02
    • 工具调用
    • 结构化输出
    • 读图
    • 长上下文
    • 深度思考
  • claude-opus-4-8
    可直接调用
    前沿
    上下文
    1000K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥9.47·¥5.62
    输出¥47.37·¥28.08
    缓存读¥0.947·¥0.562
    缓存写¥11.84·¥7.02
    • 工具调用
    • 结构化输出
    • 读图
    • 长上下文
    • 深度思考
  • claude-sonnet-5
    可直接调用
    均衡
    上下文
    1000K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥3.79·¥2.25
    输出¥18.95·¥11.23
    缓存读¥0.379·¥0.225
    缓存写¥4.74·¥2.81
    • 工具调用
    • 结构化输出
    • 读图
    • 长上下文
  • claude-fable-5-1
    可直接调用
    前沿
    上下文
    1000K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥67.37·¥39.93
    输出¥336.84·¥199.65
    缓存读¥1.68·¥0.998
    缓存写¥84.21·¥49.91
    • 工具调用
    • 结构化输出
    • 读图
    • 长上下文
    • 深度思考
  • grok-4.5
    可直接调用
    均衡
    上下文
    500K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥3.40·¥2.02
    输出¥10.20·¥6.05
    缓存读¥0.51·¥0.302
    缓存写¥0.20·¥0.119
    • 工具调用
    • 结构化输出
    • 长上下文
  • glm-5.3-flash
    可直接调用
    国产
    上下文
    128K
    大陆
    可用
    每百万 token后付费 · 套餐(pro5)
    输入¥0.08·¥0.0474
    输出¥0.28·¥0.166
    缓存读¥0.023·¥0.0136
    缓存写¥0.10·¥0.0593
    • 工具调用
    • 结构化输出
    • 深度思考
  • grok-4.6
    可直接调用
    均衡
    上下文
    500K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥3.40·¥2.02
    输出¥10.20·¥6.05
    缓存读¥0.85·¥0.504
    缓存写¥0.20·¥0.119
    • 工具调用
    • 结构化输出
    • 长上下文
  • gpt-5.6
    可直接调用
    前沿
    上下文
    922K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥3.00·¥1.78
    输出¥14.99·¥8.88
    缓存读¥0.30·¥0.178
    缓存写¥3.75·¥2.22
    • 工具调用
    • 结构化输出
    • 读图
    • 长上下文
    • responses
  • gpt-5.5
    可直接调用
    前沿
    上下文
    1050K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥3.75·¥2.22
    输出¥22.48·¥13.33
    缓存读¥0.375·¥0.222
    缓存写¥0.50·¥0.296
    • 工具调用
    • 结构化输出
    • 读图
    • 长上下文
    • responses
  • gpt-6-astra
    可直接调用
    前沿
    上下文
    922K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥7.49·¥4.44
    输出¥37.47·¥22.21
    缓存读¥0.749·¥0.444
    缓存写¥9.37·¥5.55
    • 工具调用
    • 结构化输出
    • 读图
    • 长上下文
    • responses
  • gpt-image-2
    暂不可用
    均衡
    上下文
    1K
    大陆
    不提供
    每百万 token后付费 · 套餐(pro5)
    输入¥0.16·¥0.0948
    输出¥0.84·¥0.498
    缓存读¥0.016·¥0.00948
    缓存写¥0.20·¥0.119
    • image

Every available model

Models within a tier can back each other up. Name a specific model, or name only the tier and let Prism choose inside it.

ModelTierInputOutputContextCapabilitiesRegionStatus
frontier-aFrontier0.201.001050KTool useVisionStructured outputLong contextOutside CN onlyOperational
frontier-bFrontier0.201.001050KTool useVisionStructured outputLong contextOutside CN onlyOperational
balanced-aBalanced0.080.421000KTool useVisionStructured outputLong contextOutside CN onlyOperational
balanced-bBalanced0.080.421000KTool useVisionStructured outputLong contextOutside CN onlyOperational
balanced-cBalanced0.080.421000KTool useVisionStructured outputLong contextOutside CN onlyOperational
domestic-aDomestic flagship0.030.18128KTool useStructured outputAvailable in CNOperational
domestic-bDomestic flagship0.030.18128KTool useStructured outputAvailable in CNOperational
domestic-cDomestic flagship0.030.18128KTool useStructured outputAvailable in CNOperational
fast-aFast & light0.010.060K—Outside CN onlyOperational
fast-bFast & light0.010.060K—Outside CN onlyOperational
fast-cFast & light0.010.060K—Outside CN onlyOperational
fast-dFast & light0.010.060K—Outside CN onlyOperational

Status comes from real probes every 15 minutes, not a hand-maintained table. The full uptime history lives on the status page, retained 90 days.

Model identifiers are placeholders. At launch this table is driven live by the catalogue service with real model names and current coefficients.

Which tier for which job

We'd rather you didn't overspend. This table is ordered by which tier is cheapest for a given result — not by which tier is most expensive.

What you're doingSuggestedWhy
Long agent chains, cross-file refactorsFrontierOne wrong step compounds. Redoing the work costs far more than the money you'd save.
Everyday coding and reviewBalancedThe default for most people. On this kind of work the quality gap is essentially imperceptible.
Long-form Chinese, research digestsDomestic flagshipGenuinely stronger on Chinese material, and cheaper as a side effect. This is selection, not downgrading.
Bulk classification, format conversionFast & lightThe task has no difficulty in it. Paying frontier rates here is pure waste.
Completion, inline suggestionsFast & lightLatency matters more than quality here. The fast one is the correct one.
Not sureLet Prism decideauto judges each request on its actual difficulty, which beats picking one tier up front.

Multiplier change history

Publishing only the current value leaves you guessing whether we've changed it. So here's the whole record: every change, when it was announced, when it took effect.

No changes yet — the service is not live, so this table is empty. It is empty because nothing has happened, not because nothing was recorded; the first change gets appended here with both its announcement and effective dates.

A multiplier **increase** is announced 30 days ahead and applies only to new periods — a subscription you already hold is unaffected within its period. Decreases (cheaper for you) take effect immediately.

Regional availability

Models marked "outside CN only" have not completed generative-AI service filing in mainland China and, per regulation, are not offered to users there.

  1. 01

    The console filters by your location

    Unavailable models are greyed out with the reason shown — you won't pick one and then have it fail. Prism's routing excludes them automatically too.

  2. 02

    Please state your location accurately

    Circumventing regional restrictions to use a model not offered where you are is the user's own responsibility. This is section 4 of the terms of service.

  3. 03

    Regional rules change

    Filing status is not static. When a model's availability changes we notify you in advance and let you switch tiers, or refund the unused portion.

You've seen the rates — now see what saves you money

The same quota, routed by Prism according to task difficulty, typically gets more than twice the work done. Rates are the raw material; routing is where the saving happens.