Three steps, five minutes
No credit card, no currency exchange, no documentation to read. It works even if this is your first time using a model service.
Create an account
An email address is enough. No payment required to sign up, and no card on file.
Run one line
Create a key in the console, then run the line below. The script asks you to paste the key and writes the endpoint and key to ~/.config/agiplan/env. It installs nothing and leaves your shell config alone.
Open the tool you already use
Just work. To change models, change the model field in your request: a specific model pins it, auto hands it to Prism.
# Step two is this one line curl -fsSL https://agiplan.dev/setup -o setup.sh && sh setup.sh # Load it into this shell, then check the connection . ~/.config/agiplan/env curl -s "$API_BASE_URL/models" -H "Authorization: Bearer $API_KEY"
Everything the script does is documented, and you can skip it entirely — two environment variables do the same job. The key goes into a config file in your home directory (mode 600), outside any project, so it can't be committed to git by accident. The command-line tool isn't released yet, and the script doesn't pretend to install anything.
Signing up costs nothing. Pick a plan and create a key in the console once you're in.
Whatever you use today probably just works
We speak both mainstream API protocols. If a tool lets you change its endpoint, changing that one setting is the whole integration. No plugin, no code changes.
Don't bet your whole workflow
on a single model.
Models get replaced every quarter and prices halve every six months. Whatever is best today may not be around next year. We sort models into four capability tiers; one quota drains at each tier's own rate, and switching is a single click.
Most people assume domestic models are the budget fallback. For long-form Chinese and bulk processing they are simply the better fit — and cheaper as a side effect.
| Capability tier | Input | Output | When it's the smart choice |
|---|---|---|---|
| Frontier | 0.20 | 1.00 | Long agent chains, complex refactors, work that has to be right the first time |
| Balanced | 0.08 | 0.42 | Everyday coding and review — the default for most people |
| Domestic flagship | 0.03 | 0.18 | Long-form Chinese, research digests, cost-sensitive batch work |
| Fast & light | 0.01 | 0.06 | Completion, classification, formatting — anything latency-bound |
This table is published at /models, and any multiplier increase is announced 30 days in advance. The same 480,000 CU spent entirely on the frontier tier, versus routing everyday work to Balanced and bulk work to Domestic, differ by several times over. How you spend it is your call. Our job is to keep the math honest.
Prism sends each task where it belongs
The work you do in a day varies enormously in difficulty. Renaming a variable and refactoring a module run on the same expensive model — that's exactly how the money burns.
Prism is our routing layer. Every request passes through it first, gets assessed for difficulty and required capabilities, then lands at the right point on the spectrum. Use the auto preset we maintain, or write your own rules.
See how much more the same money does
Computed from the published multipliersThis is an estimate based on task mix, not a guarantee. Your real multiplier depends on your actual distribution — the console sends you a measured comparison every month.
Want to pick the model yourself? Always an option
One model field, two ways to use it: name a specific model and it's pinned to that one; put auto or your own custom model name there and Prism routes it. Switching between the two changes nothing else — no reconfiguration, no code churn.
model: "model-a" // Pinned to this model, Prism stays out of it model: "auto" // Handed to Prism, picked by task difficulty model: "coder" // Your own custom model, your own rules
A model name of your own that picks a model per turn
Answer four questions and you get a model name of your own, such as coder. Call it and taste-dependent or hard work goes to a strong model, clearly simple work to a cheap one, and the rest to the default you chose. Every rule and threshold is laid out in the console; replay a change against your own traffic, then save it as a new version.
Drag it and see who gets each message
Sample conversation · illustrative probabilities · same rule code as productionyouMake the login page buttons and spacing look nicer
Frontend0.84Taste0.78Hard0.10Strong modelFrontierDoing it well needs tasteyouWrite a script that unpacks the .gz files in logs/ and files them by date
Script0.90Taste0.03Hard0.06Cheap modelFast & lightClearly simple workyouWhat does this error mean: Cannot read properties of undefined
Debug0.78Taste0.02Hard0.16Cheap modelFast & lightClearly simple workyouMove the order module from callbacks to async/await without changing behavior
Refactor0.80Taste0.04Hard0.58Strong modelFrontierJudged hardyouAdd a pagination field to this endpoint without breaking old clients
Backend0.71Taste0.05Hard0.30Default modelBalancedNone of the above, defaultyouWrite a README section for this endpoint
Writing0.74Taste0.28Hard0.05Default modelBalancedNone of the above, defaultyouDesign data isolation for multi-tenancy and list the trade-offs
Architecture0.86Taste0.12Hard0.72Strong modelFrontierJudged hardyouHi, quick question first
Chat0.95Taste0.01Hard0.01Cheap modelFast & lightClearly simple work
Answer four questions
What you mostly use it for, how to trade quality against cost, who fills the strong, default and cheap positions, and whether to also create quality- and cost-leaning versions for subagents. The rules come from those answers; nothing is hidden.
Replay your own traffic
Change a threshold and immediately see how the last 30 days of requests would split, how much more or less work the quota does and which requests would move. It runs in your browser, and every saved version can be restored in one click.
Fill your tool's three slots
Claude Code already splits the main task and subagents into three model slots. Put coder-high, coder and coder-lite there: the tool picks the slot, we pick the model within it by task.
Judging looks at your latest message and discards it; we keep only numbers such as the judged probabilities, never conversation content. Saving can be turned off and cleared at any time.
Routing by task type is planned and not yet open. You can already create and call a custom model; until it opens, every turn goes to the model you set for “when unjudged”, and nothing needs to change on your side when it does.
Cheaper must never mean worse. When difficulty lands near a boundary, auto always rounds up — paying a little more beats degraded output every time, and that is hard-coded. In a custom model the lean is yours: work only goes down as far as you tilt it toward cost, and every such step is recorded in the routing trace.
If a model fails, we switch to a same-tier backup
Any model can hit a rate limit, time out, or go down. Prism immediately hands the request to a backup model in the same capability tier, so this call doesn't fail. It's on by default, and you can turn it off.
Turn it off if you prefer
On by default. With `"failover": { "enabled": false }` in a custom model, a model failure returns an error instead of switching. Naming your own backup order is not built yet: selection is automatic within the tier, and `scope` accepts only `same_tier`.
We only switch within a tier
We switch only within the tier you chose and never across tiers to a cheaper one — for that kind of downgrade we would rather fail and tell you (503, with the full record of attempts) than swap quietly. Within the tier we prefer the higher-quality route, so a failure can land you on a slightly lower-ranked one; every response carries `x-agiplan-model-actual`, so which model answered is never hidden.
Failed attempts cost nothing
Only the attempt that actually returns is metered. Whatever we tried in between never reaches your bill.
Whichever line serves you is the one you pay for
After a failover you are billed at the price of the line that actually served the request, and the bill names that model. The attempts that failed never reach your bill — you shouldn't pay for our outage.
Every routing decision can be replayed
Automatic routing naturally invites the suspicion that you're being quietly served something cheap. The only answer is to show you every decision. Each request keeps a trace, retained 30 days, exportable.
- 0msauthap_live_…f3c2 · 5× Pro
- 3msquotareserved 11.742 CU (14,200 in + max_tokens 8,192, worst case), all three windows clear
- 11msprismdifficulty 0.82 · tools=3 tokens_in=14,200 code_blocks=2
- 12msrouterule #1 matched → frontier · model-a
- 14msrespondmodel-a rate limited
- 210msfailoversame-tier backup model-b · billed at model-b's rate
- 224msstreamfirst token 890ms
- 8.42scommitin 14,200 / out 2,130 → 4.970 CU billed, remaining 6.772 CU of the reservation released
Traces hold metadata only, never request bodies — the same architecture as the zero-retention promise in section 07. When something breaks, send us the trace_id and support can pull it up instantly.
How quota is counted, out in the open
Price matters less than being able to check it. We use one unit, the CU: 1 CU = 1,000 output tokens on the reference model
Every model's input and output coefficients are published, and every call is inspectable, exportable and reconcilable. For a concrete sense of scale, here's what each plan is worth:
| Plan | Monthly quota | Roughly equivalent to |
|---|---|---|
| Lite | 30,000 CU | An entry-level monthly plan |
| Pro | 96,000 CU | A standard subscription tier |
| 5× Pro | 480,000 CU | 5× a standard subscription |
| Max | 1,920,000 CU | 20× a standard subscription |
5 hours · week · billing month
All three apply at once and are live in the console. The billing month is anchored to your signup date, not the calendar — subscribing on the 29th doesn't leave you with two days of quota.
Every single call is inspectable
Timestamp, model, input and output tokens, CU charged, latency, trace ID. Exportable as CSV by project and date range.
Increases announced 30 days ahead
Terms for a subscription already sold never change mid-period. Plan versions are immutable; a price change applies only to new subscriptions.
Refunded pro rata
Cancel any time and get the unused portion of the period back. We don't make money on quota you never spend.
Whether each model is up right now
The status page gives you, per model: whether it is up right now, uptime over the last 30 days, whether it is available in your region, and every raw probe record (timestamp, outcome, latency, failure kind — exportable by month). Driven by real probes every 15 minutes. Multipliers live on /models. Latency is internal only for now; we do not publish it.
We tell you before you run out
At 80% of your weekly or monthly quota (the threshold is fixed for now, not configurable) you get a heads-up plus a recommendation: which tier to switch to and how much longer it buys you, or what upgrading would cost. The 5-hour window sends no email — it resets several times a day, so mail would be noise; that figure is in every response header and on the live ring in your console.
We don't compete on price. That effort goes somewhere else: staying up, adding up, and owning it when things break.
Enough to try it properly
- 8,000 CU / week
- 30,000 CU / month
- 5-hour window 600 CU
- Choose from every model
- Prism auto routing
≈ a standard subscription
- 24,000 CU / week
- 96,000 CU / month
- 5-hour window 1,800 CU
- Choose from every model
- Email support, answered by a person
≈ 5× Pro
- 120,000 CU / week
- 480,000 CU / month
- 5-hour window 9,000 CU
- Custom Prism models
- Priority routing (planned)
≈ 20× Pro
- 480,000 CU / week
- 1,920,000 CU / month
- 5-hour window 30,000 CU
- 8 concurrent requests
- Dedicated routing (planned)
One contract, your whole team's usage and cost under control.
For teams the hard part usually isn't access — it's reconciling the bill at month end. All of this runs on our own gateway: quota, routing, metering and audit are one codebase, not middleware bolted together.
Central purchasing and billing
Buy by seat or by quota pack, invoiced once, paid by transfer. Annual master agreements and prepayment supported.
Department and project quotas
Split the pool across departments, projects and individuals, each with its own weekly and monthly ceilings and alert thresholds. One team overrunning never affects another.
Cost attribution
Consumption broken out by project, member and model, exportable to your own BI. Who spent what, at a glance.
SSO and member governancePlanned
Planned: OIDC and SAML plus the major workplace identity providers, joiner-leaver sync. Key scoping and forced rotation work today; single sign-on is not built yet.
A shared Prism policy
Define one routing policy for the team and push it to everyone. Costs stay predictable and nobody has to guess which model to use.
Private deployment and compliancePlanned
DPAs, security questionnaires and audit support are available today. Gateway nodes inside your own VPC are not available yet — talk to us first rather than planning around them.
Your code and prompts are not stored by default
You're about to send us the most sensitive material you have — source code, internal documents, client data. This section is wordier than a marketing page should be. Please actually read it.
Request and response bodies never hit disk
The gateway records only what metering requires: model, token counts, latency, status code, trace ID. Bodies are forwarded in memory and released the moment the stream ends — no logs, no database, no object storage.
We haven't built this switch
Seeing the body would obviously help when debugging, and we have not built it — bodies are never written to disk, with no exception. If we ever do build it, it will be off by default, enabled only by you, bounded by a stated retention period, and announced in advance. Until then there is no path at all by which we could see your bodies.
No training, no resale, no human reading
Your content never enters any training or fine-tuning pipeline, is never handed to third parties, and staff have no access. All three are written into the terms of service and legally binding.
Encrypted end to end
TLS 1.3 externally, mutual TLS between internal services, and encryption enforced on upstream connections too.
Hashed — we can't read them either
API keys are stored hashed and shown in plaintext exactly once, at creation — we can't read them back either, so a lost key is rotated rather than recovered.
Least privilege, fully audited
Every back-office action touching an account is recorded: who, when, what changed, from which IP. Audit logs are tamper-evident and can be produced on request.
Export or delete at any time
Usage data is exportable whenever you want. On closure, your email, credentials, keys and technical logs are destroyed within 30 days; the transaction and usage ledgers cannot be deleted and are instead unlinked from you — those three record types are append-only, and a ledger that can be edited proves nothing in a billing dispute. The privacy policy spells this out. (Account closure is not built yet; see the terms below.)
You hear from us within 72 hours
If a security incident could affect your data, we notify affected users within 72 hours and publish a post-mortem.
How we make money, stated plainly
The most reliable way to predict whether a company will treat you badly is to understand how it earns. So here it is.
We earn the spread on purchasing scale and routing efficiency.
We buy at volumes far beyond any individual user and turn that scale, plus Prism's routing efficiency, into your price. This business only works if you stay: acquiring a customer costs far more than a month of gross margin, so burning you once is a net loss.
We don't earn on quota you never spend — cancellations are refunded pro rata. We don't earn by giving you less than you asked for — the model you pick is the model that runs. We don't earn on your data — no training, no resale.
Long term the profit comes from teams and enterprises, and enterprise buyers only work with vendors that have a clean reputation. So there's no goodwill involved here. We just did the math.
Eight things we will never do
This list becomes a standalone Service Commitments page and takes effect alongside the terms. Break any one of them and you can demand a full refund.
- 01
Never silently substitute the model you chose
No reduced parameters, no truncated context, no quiet downgrade at peak hours. The response header carries the model that actually ran, so you can check for yourself.
- 02
Never train on, resell, or read your content
Your prompts and outputs enter no training pipeline, go to no third party, and staff have no access.
- 03
Never reduce quota terms mid-subscription
Any change to multipliers or quotas applies to new periods only; increases are announced 30 days in advance.
- 04
Never use countdowns, scarcity, or referral bounties
The price is the price. A countdown exists to stop you thinking it over, and we don't need that conversion.
- 05
Never claim official authorisation from any provider
We are an independent aggregation gateway and will not say otherwise. When another platform does, ask them for it in writing.
- 06
Never pretend an outage didn't happen
Any incident affecting more than 1% of users gets a signed post-mortem within 48 hours: what happened, root cause, what we changed.
- 07
Never put obstacles in front of a refund
Refund rules are fixed in the terms and self-service in the console. No support ticket, no reason required, no retention script.
- 08
Never disappear without notice
If we ever shut down, you get 60 days' notice, a full refund of unused quota, and help migrating elsewhere.
Your request bodies have a market price
This rarely gets said out loud. Every prompt and every file a user sends can be recorded in full — distilled into smaller models, packaged and sold as training data, or quietly traded between platforms. No notification is sent, and the terms usually contain no explicit prohibition.
The internal documents, client data and unreleased code you pasted in may already be in somebody's training set.
We're against this, and not only in words.
- 01
Zero retention is architecture, not policy
There is no code path in the gateway that writes a request body anywhere. Not "we promise not to look" — there is nowhere to put it. This is auditable.
- 02
We buy no data of unclear origin
We purchase request corpora from no platform and no data broker, at any price, under any label.
- 03
We refuse data-for-price deals
When a provider offers a lower rate in exchange for data sharing, we decline — even though it means higher costs and thinner margins for us.
- 04
We're willing to pay for it
Every line above reduces what we can earn. That's the price we chose to pay, and the reason we'd like your money instead of somebody else's.
- 05
We accept verification
All of this goes into the terms of service with legal force. We accept third-party security audits, and you're welcome to inspect the traffic yourself.
We can't change the whole industry. But whatever passes through our pipe, we can guarantee it goes nowhere else.
This is not an anonymous site.
Most services like this are anonymous: no legal entity, no names, a domain that can change overnight. Our names, engineering notes and post-mortems are all here for you to check.
You're probably wondering.
I've never used anything like this. Can I set it up myself?
Yes. Sign up, create a key in the console, and run one line that writes it into your local config. Then open the coding tool you already use — no config files to edit by hand. If you get stuck, open a ticket in the console and a person answers. We email you when a ticket gets a reply (the email carries no reply text — the content stays on the site).
What if you shut down in six months?
That's the reasonable worry, and a promise alone is worth nothing. What we can do is make leaving expensive for us: legal entity and registration published in the footer, invoices, contracts, named team. Service Commitment 8 states that if we stop, you get 60 days' notice, a full refund of unused quota and migration help. You can also start on ¥49 Lite for a month — there's no annual lock-in.
Could my quota be quietly reduced?
No. Plan versions are immutable: the quota terms and multipliers at the moment you subscribed are fixed for the whole period, and any change applies to new periods only with 7 days' notice. That's a database-level constraint, not just a promise — what's stored is the version you bought.
With Prism routing, am I quietly getting cheap models?
Yes — but only when the task doesn't need a stronger one. The logic and the outcome are both visible to you: every request's trace carries the difficulty score, the rule that matched and the model that actually ran, retained 30 days and exportable. Borderline difficulty always rounds up, and requests declaring tool use or long context exclude models that can't serve them. You can also skip Prism entirely and name the model yourself.
What happens when a model goes down? Do I pay more?
Failover is on by default: Prism moves the request to a backup model in the same capability tier, so your call doesn't fail. Failed attempts aren't billed — only the one that returns. If the only backup costs more, you're still billed at the tier you requested and we absorb the difference. You can turn failover off (set failover.enabled to false in a mix); you can't currently choose the backups or their order — Prism picks within the same tier.
What if I run out of quota early?
Two options today: switch to a cheaper capability tier and keep going; or upgrade, with the difference prorated over the remaining days and your billing date unchanged. There is a third route: pay-as-you-go, drawing from your balance at ¥1 = 500 CU with the same per-model rates as your plan. Off by default; enabling it requires a monthly ceiling, and charging stops there.
Can anyone read my code and prompts?
Not stored by default. The gateway records metering metadata only; bodies are forwarded in memory and released when the stream ends. No training, no resale, no staff access — all three are written into the terms with legal force.
Is every model available everywhere?
No. Some models have not completed generative-AI service filing in mainland China and are therefore not offered to users there; the console labels regional availability per model. Please state your location accurately — circumventing regional restrictions is the user's own responsibility, as set out in section 4 of the terms.
Why you instead of buying from the provider directly?
Three reasons: CNY billing and invoices, so no credit card or FX; one quota covering models from several providers including domestic ones, switchable without code changes; and Prism, which stretches the same money further — how much further depends on how much of your work genuinely needs the frontier tier: below roughly 30% frontier share the same quota does about twice the work; if 90% of your calls must run on frontier, the difference is small. The comparator on /prism lets you work it out yourself. If you use exactly one model from one provider and already have a working payment method, going direct is simpler. We don't think we need to convince everyone.
Are you officially authorised?
No, and we won't claim to be. We are an independent multi-provider aggregation gateway. When a comparable platform claims official authorisation, ask them to show it in writing.
Compute should work like utilities:metered transparently, drawn on demand,and switchable to another supplier at will.
A credit card, an exchange rate, a regional restriction — none of these should be the reason someone can't use the best AI available. That's why we started.
Then we found a second reason. Good models are becoming accessible, but the price is often that you hand over your data without knowing it. That price shouldn't exist, and we'll spend what we have to hold it back.
What we build is the pipe in the middle. We want it sturdy, we want the accounting honest, and we want you free to leave whenever you like. While you're using it, it's only a pipe — it doesn't keep a single word.
Try it for a month, then decide whether to trust us
No annual lock-in, cancel whenever. Full refund within 7 days if you've used under 10% — self-service in the console, no ticket required.