Prism

Every routing rule published, including the ones that cost us

Automatic routing naturally invites the suspicion that you're being quietly served something cheap. Anyone can write marketing copy; nobody can write rule details they haven't built. So this page reads like documentation.

Four layers, from fully manual to fully managed

All four share one trace format. Whichever layer handled it, any request can be replayed to show why that model was chosen.

Your requestCalled normally, no code changesPRISMFrontier · 1.00×Balanced · 0.42×Domestic flagship · 0.18×Fast & light · 0.06×
LayerWhat it doesWho it's for
Named modelPut a specific model ID in the field and Prism stays out of it entirely.People who already know what they want.
Prism autoThe default policy we maintain and keep tuning. Works the moment you install.People who don't want to configure anything. It's the default for new accounts.
Custom mixWrite your own rules, save them under a name, call it like a model.People with a predictable task distribution who want tight cost control.
FailoverThe backstop beneath all three. Moves to a same-tier backup when a model is unavailable.Everyone. On by default, switchable off; backups are picked within the same tier, not chosen by you.

Try it yourself

All four weights are adjustable and normalise as you drag. These are the real coefficients published at /models, not demo numbers.

Common distributions
How your work is distributed across the tiers
20%
48%
20%
12%
Without Prism, frontier tier throughout480,000 CU of work
Routed on this distribution, same capability1,079,137 CU of work
2.2×The same 5× Pro quota gets 599,137 CU of extra work done.

The multiplier is the weighted result of your distribution and the published coefficients — check it yourself against /models. Your real distribution is recomputed monthly by the console from actual usage.

How difficulty is judged

The classifier finishes before the request reaches an upstream, with a latency budget under 15ms. It has to be a local small model or rule scoring — spending an LLM call to save money doesn't add up.

SignalWhat it tells us
Input token countA long context usually means a more involved task.
Whether tools are declaredRequests carrying tools are usually agent work, where there's little room for error.
Conversation depthAccumulated multi-turn work is generally harder than a single turn.
Code block count and lengthSeparates "what's the syntax" from "rewrite this module".
Intent keywordsRefactor, design, debug, translate and format differ enormously in difficulty.
Difficulty earlier in the session**Not in effect yet.** The scorer has a term for it, but nothing feeds in how hard earlier turns were, so it is currently always zero — see the "escalate on retry" open decision. It is listed here because this page promises every signal, including the one that isn't wired.
Explicit request headerX-AGIPlan-Difficulty overrides the automatic judgement outright.

Version one is plain weighted rule scoring. Once live we train a small classifier on real data, using "classifier said easy but the user retried" as negative samples — that feedback is far more accurate than hand labelling.

Three hard constraints, not settings

Cheaper must never mean worse, or you'd leave in the first week. These three are fixed in code and no mix or policy can get around them.

  1. 01

    Round up, never down

    When difficulty lands near a threshold it always rounds up. The cost of being wrong is asymmetric: paying a bit more beats degraded output, and degraded output is something you may not notice at the time.

  2. 02

    Capability requirements filter before cost

    When a request declares tool use, vision or a minimum context length, models that can't serve it are excluded and never enter the cost comparison. A cheap model that can't do the job isn't an option.

  3. 03

    Compliance is not an optimisation

    Models without a mainland China filing are not served to users in mainland China. The route is excluded at admission — not relaxed because there is no alternative, and not allowed through because your own mix names it.

Failover and billing

Upstream unavailability is normal, not exceptional. On by default, confined to the same capability tier, and failed attempts are never billed.

Triggers

5xx, rate limits, timeouts, credential failure

First-token timeout defaults to 60 seconds; send x-agiplan-first-byte-timeout-ms to change it, clamped to 5–60 seconds. This applies to streaming requests only: a non-streaming upstream replies only once generation is done, so it is bounded by the 10-minute total limit instead. Truncated output that wasn't your intent is logged but never auto-retried — auto-retry would double-bill.

Scope

Same capability tier only

Only swaps within the same capability tier so output quality is unchanged. Cross-tier fallback isn't offered — not even as a config option. A mix can turn failover off (failover.enabled = false), which means fail rather than substitute.

When the tier is exhausted

Fail and say so, never silently downgrade

If no model in the tier is available, the default is a 503 with the full record of attempts — not a quietly weaker model.

Billing

Failures aren't billed; success is billed as served

Only the attempt that returns is metered — the ones we tried and dropped never reach your bill. Whichever line ends up serving you is the one you pay for, and the bill names that model. One exception, stated plainly: if the upstream dies after the stream has started, what was already produced is billed. Those tokens really were generated — same reasoning as when you abort it yourself.

"failover": {
  "enabled": true,                 // defaults to true
  "scope": "same_tier"             // same_tier only for now; custom chains are not built
}
Failover options

Routing traces

Not a bonus feature — the precondition for automatic routing being trustworthy at all. Every request keeps a trace, retained 30 days, exportable.

EXAMPLEtrace tr_01J8FQ3M7XTotal 8.42s · Billed 4.970 CU
  1. 0msauthap_live_…f3c2 · 5× Pro
  2. 3msquotareserved 11.742 CU (14,200 in + max_tokens 8,192, worst case), all three windows clear
  3. 11msprismdifficulty 0.82 · tools=3 tokens_in=14,200 code_blocks=2
  4. 12msrouterule #1 matched → frontier · model-a
  5. 14msrespondmodel-a rate limited
  6. 210msfailoversame-tier backup model-b · billed at model-b's rate
  7. 224msstreamfirst token 890ms
  8. 8.42scommitin 14,200 / out 2,130 → 4.970 CU billed, remaining 6.772 CU of the reservation released

Traces hold metadata only, never request bodies — the same architecture as zero retention. They contain model identifiers only, never any supply-side identifier.

Response header

X-AGIPlan-Trace

Successful responses carry a trace ID plus X-AGIPlan-Model-Actual naming the model that really ran. Store them in your own logs if you like. Refused responses carry no trace ID — that call left no usage record to look up; on a 503 the chain is in the body as `attempts`.

Console

Expand any usage row

Each line expands into the full timeline. Requests that went through failover are marked so you can filter them out at a glance.

API

GET /v1/traces/{id}

Fetch it with your own key and get structured JSON. Pipe it into your own observability stack if you want.

Support

Send us the trace_id

Support pulls it up by ID — you don't have to reconstruct what happened.

Example mix configuration

auto and custom models now share one rule language: rules are matched top to bottom and the first hit wins; no match falls through to default (or is refused when default is null). Name it, save it, and call it like any model.

{
  "name": "my-coding-model",
  "pools": {
    "strong":   { "tier": "frontier" },
    "domestic": { "tier": "domestic" },
    "mid":      { "tier": "balanced" },
    "cheap":    { "tier": "fast" }
  },
  "rules": [
    { "id": "hard",    "when": { "difficulty": { "gte": 0.75 } },                        "use": "strong" },
    { "id": "zh_long", "when": { "lang": "zh", "tokens_in": { "gt": 8000 } },             "use": "domestic" },
    { "id": "medium",  "when": { "difficulty": { "gte": 0.35, "lt": 0.75 } },             "use": "mid" },
    { "id": "simple",  "when": { "difficulty": { "lt": 0.35 } },                          "use": "cheap" }
  ],
  "default": "mid",
  "constraints": {
    "require_tools": true,       // models without tool use are excluded outright
    "min_context": 128000,
    "max_cu_per_request": 500    // reject anything larger, so one request can't drain the quota
  },
  "failover": { "enabled": true, "scope": "same_tier" }
}
A complete custom-model policy

Rules read — now work out what it saves

Rates are the raw material; routing is where the saving happens. The pricing page has an advisor that recommends a plan from how you actually work.