Prism
Every routing rule published, including the ones that cost us
Automatic routing naturally invites the suspicion that you're being quietly served something cheap. Anyone can write marketing copy; nobody can write rule details they haven't built. So this page reads like documentation.
Four layers, from fully manual to fully managed
All four share one trace format. Whichever layer handled it, any request can be replayed to show why that model was chosen.
| Layer | What it does | Who it's for |
|---|---|---|
| Named model | Put a specific model ID in the field and Prism stays out of it entirely. | People who already know what they want. |
| Prism auto | The default policy we maintain and keep tuning. Works the moment you install. | People who don't want to configure anything. It's the default for new accounts. |
| Custom mix | Write your own rules, save them under a name, call it like a model. | People with a predictable task distribution who want tight cost control. |
| Failover | The backstop beneath all three. Moves to a same-tier backup when a model is unavailable. | Everyone. On by default, switchable off; backups are picked within the same tier, not chosen by you. |
Try it yourself
All four weights are adjustable and normalise as you drag. These are the real coefficients published at /models, not demo numbers.
The multiplier is the weighted result of your distribution and the published coefficients — check it yourself against /models. Your real distribution is recomputed monthly by the console from actual usage.
How difficulty is judged
The classifier finishes before the request reaches an upstream, with a latency budget under 15ms. It has to be a local small model or rule scoring — spending an LLM call to save money doesn't add up.
| Signal | What it tells us |
|---|---|
| Input token count | A long context usually means a more involved task. |
| Whether tools are declared | Requests carrying tools are usually agent work, where there's little room for error. |
| Conversation depth | Accumulated multi-turn work is generally harder than a single turn. |
| Code block count and length | Separates "what's the syntax" from "rewrite this module". |
| Intent keywords | Refactor, design, debug, translate and format differ enormously in difficulty. |
| Difficulty earlier in the session | **Not in effect yet.** The scorer has a term for it, but nothing feeds in how hard earlier turns were, so it is currently always zero — see the "escalate on retry" open decision. It is listed here because this page promises every signal, including the one that isn't wired. |
| Explicit request header | X-AGIPlan-Difficulty overrides the automatic judgement outright. |
Version one is plain weighted rule scoring. Once live we train a small classifier on real data, using "classifier said easy but the user retried" as negative samples — that feedback is far more accurate than hand labelling.
Three hard constraints, not settings
Cheaper must never mean worse, or you'd leave in the first week. These three are fixed in code and no mix or policy can get around them.
- 01
Round up, never down
When difficulty lands near a threshold it always rounds up. The cost of being wrong is asymmetric: paying a bit more beats degraded output, and degraded output is something you may not notice at the time.
- 02
Capability requirements filter before cost
When a request declares tool use, vision or a minimum context length, models that can't serve it are excluded and never enter the cost comparison. A cheap model that can't do the job isn't an option.
- 03
Compliance is not an optimisation
Models without a mainland China filing are not served to users in mainland China. The route is excluded at admission — not relaxed because there is no alternative, and not allowed through because your own mix names it.
Failover and billing
Upstream unavailability is normal, not exceptional. On by default, confined to the same capability tier, and failed attempts are never billed.
5xx, rate limits, timeouts, credential failure
First-token timeout defaults to 60 seconds; send x-agiplan-first-byte-timeout-ms to change it, clamped to 5–60 seconds. This applies to streaming requests only: a non-streaming upstream replies only once generation is done, so it is bounded by the 10-minute total limit instead. Truncated output that wasn't your intent is logged but never auto-retried — auto-retry would double-bill.
Same capability tier only
Only swaps within the same capability tier so output quality is unchanged. Cross-tier fallback isn't offered — not even as a config option. A mix can turn failover off (failover.enabled = false), which means fail rather than substitute.
Fail and say so, never silently downgrade
If no model in the tier is available, the default is a 503 with the full record of attempts — not a quietly weaker model.
Failures aren't billed; success is billed as served
Only the attempt that returns is metered — the ones we tried and dropped never reach your bill. Whichever line ends up serving you is the one you pay for, and the bill names that model. One exception, stated plainly: if the upstream dies after the stream has started, what was already produced is billed. Those tokens really were generated — same reasoning as when you abort it yourself.
"failover": { "enabled": true, // defaults to true "scope": "same_tier" // same_tier only for now; custom chains are not built }
Routing traces
Not a bonus feature — the precondition for automatic routing being trustworthy at all. Every request keeps a trace, retained 30 days, exportable.
- 0msauthap_live_…f3c2 · 5× Pro
- 3msquotareserved 11.742 CU (14,200 in + max_tokens 8,192, worst case), all three windows clear
- 11msprismdifficulty 0.82 · tools=3 tokens_in=14,200 code_blocks=2
- 12msrouterule #1 matched → frontier · model-a
- 14msrespondmodel-a rate limited
- 210msfailoversame-tier backup model-b · billed at model-b's rate
- 224msstreamfirst token 890ms
- 8.42scommitin 14,200 / out 2,130 → 4.970 CU billed, remaining 6.772 CU of the reservation released
Traces hold metadata only, never request bodies — the same architecture as zero retention. They contain model identifiers only, never any supply-side identifier.
X-AGIPlan-Trace
Successful responses carry a trace ID plus X-AGIPlan-Model-Actual naming the model that really ran. Store them in your own logs if you like. Refused responses carry no trace ID — that call left no usage record to look up; on a 503 the chain is in the body as `attempts`.
Expand any usage row
Each line expands into the full timeline. Requests that went through failover are marked so you can filter them out at a glance.
GET /v1/traces/{id}
Fetch it with your own key and get structured JSON. Pipe it into your own observability stack if you want.
Send us the trace_id
Support pulls it up by ID — you don't have to reconstruct what happened.
Example mix configuration
auto and custom models now share one rule language: rules are matched top to bottom and the first hit wins; no match falls through to default (or is refused when default is null). Name it, save it, and call it like any model.
{ "name": "my-coding-model", "pools": { "strong": { "tier": "frontier" }, "domestic": { "tier": "domestic" }, "mid": { "tier": "balanced" }, "cheap": { "tier": "fast" } }, "rules": [ { "id": "hard", "when": { "difficulty": { "gte": 0.75 } }, "use": "strong" }, { "id": "zh_long", "when": { "lang": "zh", "tokens_in": { "gt": 8000 } }, "use": "domestic" }, { "id": "medium", "when": { "difficulty": { "gte": 0.35, "lt": 0.75 } }, "use": "mid" }, { "id": "simple", "when": { "difficulty": { "lt": 0.35 } }, "use": "cheap" } ], "default": "mid", "constraints": { "require_tools": true, // models without tool use are excluded outright "min_context": 128000, "max_cu_per_request": 500 // reject anything larger, so one request can't drain the quota }, "failover": { "enabled": true, "scope": "same_tier" } }
Rules read — now work out what it saves
Rates are the raw material; routing is where the saving happens. The pricing page has an advisor that recommends a plan from how you actually work.