One endpoint, every model

One subscription,
the whole model spectrum.

Keep the coding tools you already use. Sign up, run one line, and pick freely between frontier and domestic models.

Billed in CNY · Bilingual · No credit card · Invoices and contracts

12 models availableUptime and post-mortems are publicStatus page

Your quota, always visible

5× Pro
Billing month
39%
This week
65%
5-hour window
68%
Current model · switch any time
Frontier1.00×Balanced0.42×Domestic flagship0.18×Fast & light0.06×
01Get started

Three steps, five minutes

No credit card, no currency exchange, no documentation to read. It works even if this is your first time using a model service.

Step one

Create an account

An email address is enough. No payment required to sign up, and no card on file.

Step two

Run one line

Create a key in the console, then run the line below. The script asks you to paste the key and writes the endpoint and key to ~/.config/agiplan/env. It installs nothing and leaves your shell config alone.

Step three

Open the tool you already use

Just work. To change models, change the model field in your request: a specific model pins it, auto hands it to Prism.

shell
# Step two is this one line
curl -fsSL https://agiplan.dev/setup -o setup.sh && sh setup.sh

# Load it into this shell, then check the connection
. ~/.config/agiplan/env
curl -s "$API_BASE_URL/models" -H "Authorization: Bearer $API_KEY"

Everything the script does is documented, and you can skip it entirely — two environment variables do the same job. The key goes into a config file in your home directory (mode 600), outside any project, so it can't be committed to git by accident. The command-line tool isn't released yet, and the script doesn't pretend to install anything.

Create an account

Signing up costs nothing. Pick a plan and create a key in the console once you're in.

Whatever you use today probably just works

We speak both mainstream API protocols. If a tool lets you change its endpoint, changing that one setting is the whole integration. No plugin, no code changes.

Command-line coding assistantsIDE and editor extensionsAgent and workflow frameworksYour own apps and official SDKsAutomation scripts of every kind
02What you get

Don't bet your whole workflow
on a single model.

Models get replaced every quarter and prices halve every six months. Whatever is best today may not be around next year. We sort models into four capability tiers; one quota drains at each tier's own rate, and switching is a single click.

Most people assume domestic models are the budget fallback. For long-form Chinese and bulk processing they are simply the better fit — and cheaper as a side effect.

Capability tierInputOutputWhen it's the smart choice
Frontier0.201.00Long agent chains, complex refactors, work that has to be right the first time
Balanced0.080.42Everyday coding and review — the default for most people
Domestic flagship0.030.18Long-form Chinese, research digests, cost-sensitive batch work
Fast & light0.010.06Completion, classification, formatting — anything latency-bound

This table is published at /models, and any multiplier increase is announced 30 days in advance. The same 480,000 CU spent entirely on the frontier tier, versus routing everyday work to Balanced and bulk work to Domestic, differ by several times over. How you spend it is your call. Our job is to keep the math honest.

03Prism

Prism sends each task where it belongs

The work you do in a day varies enormously in difficulty. Renaming a variable and refactoring a module run on the same expensive model — that's exactly how the money burns.

Prism is our routing layer. Every request passes through it first, gets assessed for difficulty and required capabilities, then lands at the right point on the spectrum. Use the auto preset we maintain, or write your own rules.

Your requestCalled normally, no code changesPRISMFrontier · 1.00×Balanced · 0.42×Domestic flagship · 0.18×Fast & light · 0.06×

See how much more the same money does

Computed from the published multipliers
20%
Frontier 20%Balanced 48%Domestic flagship 20%Fast & light 12%
Without Prism, frontier tier throughout480,000 CU of work
Routed by Prism, same capability1,079,137 CU of work
2.2×The same 5× Pro quota gets 599,137 CU of extra work done — yours for nothing.
The multiplier holds at every plan. The higher the plan, the larger the absolute gain.

This is an estimate based on task mix, not a guarantee. Your real multiplier depends on your actual distribution — the console sends you a measured comparison every month.

Want to pick the model yourself? Always an option

One model field, two ways to use it: name a specific model and it's pinned to that one; put auto or your own custom model name there and Prism routes it. Switching between the two changes nothing else — no reconfiguration, no code churn.

request
model: "model-a"        // Pinned to this model, Prism stays out of it
model: "auto"           // Handed to Prism, picked by task difficulty
model: "coder"          // Your own custom model, your own rules

A model name of your own that picks a model per turn

Answer four questions and you get a model name of your own, such as coder. Call it and taste-dependent or hard work goes to a strong model, clearly simple work to a cheap one, and the rest to the default you chose. Every rule and threshold is laid out in the console; replay a change against your own traffic, then save it as a new version.

Drag it and see who gets each message

Sample conversation · illustrative probabilities · same rule code as production
Quality first: when unsure, use the strong modelCost first: send work down more readily
  1. youMake the login page buttons and spacing look nicer

    Frontend0.84Taste0.78Hard0.10
    Strong modelFrontierDoing it well needs taste
  2. youWrite a script that unpacks the .gz files in logs/ and files them by date

    Script0.90Taste0.03Hard0.06
    Cheap modelFast & lightClearly simple work
  3. youWhat does this error mean: Cannot read properties of undefined

    Debug0.78Taste0.02Hard0.16
    Cheap modelFast & lightClearly simple work
  4. youMove the order module from callbacks to async/await without changing behavior

    Refactor0.80Taste0.04Hard0.58
    Strong modelFrontierJudged hard
  5. youAdd a pagination field to this endpoint without breaking old clients

    Backend0.71Taste0.05Hard0.30
    Default modelBalancedNone of the above, default
  6. youWrite a README section for this endpoint

    Writing0.74Taste0.28Hard0.05
    Default modelBalancedNone of the above, default
  7. youDesign data isolation for multi-tenancy and list the trade-offs

    Architecture0.86Taste0.12Hard0.72
    Strong modelFrontierJudged hard
  8. youHi, quick question first

    Chat0.95Taste0.01Hard0.01
    Cheap modelFast & lightClearly simple work
These 8 turns: 3 strong · 2 default · 3 cheapThe same quota does 2.0× the work of using the frontier tier throughout
Estimated with the public tier rates, 12K context and 900 output tokens per turn, each turn on its own without conversation caching. Your real multiple depends on your task mix; once you create one, the console replays your own traffic.
1

Answer four questions

What you mostly use it for, how to trade quality against cost, who fills the strong, default and cheap positions, and whether to also create quality- and cost-leaning versions for subagents. The rules come from those answers; nothing is hidden.

2

Replay your own traffic

Change a threshold and immediately see how the last 30 days of requests would split, how much more or less work the quota does and which requests would move. It runs in your browser, and every saved version can be restored in one click.

3

Fill your tool's three slots

Claude Code already splits the main task and subagents into three model slots. Put coder-high, coder and coder-lite there: the tool picks the slot, we pick the model within it by task.

Create a custom model

Judging looks at your latest message and discards it; we keep only numbers such as the judged probabilities, never conversation content. Saving can be turned off and cleared at any time.

Routing by task type is planned and not yet open. You can already create and call a custom model; until it opens, every turn goes to the model you set for “when unjudged”, and nothing needs to change on your side when it does.

Cheaper must never mean worse. When difficulty lands near a boundary, auto always rounds up — paying a little more beats degraded output every time, and that is hard-coded. In a custom model the lean is yours: work only goes down as far as you tilt it toward cost, and every such step is recorded in the routing trace.

If a model fails, we switch to a same-tier backup

Any model can hit a rate limit, time out, or go down. Prism immediately hands the request to a backup model in the same capability tier, so this call doesn't fail. It's on by default, and you can turn it off.

Configurable

Turn it off if you prefer

On by default. With `"failover": { "enabled": false }` in a custom model, a model failure returns an error instead of switching. Naming your own backup order is not built yet: selection is automatic within the tier, and `scope` accepts only `same_tier`.

Same tier

We only switch within a tier

We switch only within the tier you chose and never across tiers to a cheaper one — for that kind of downgrade we would rather fail and tell you (503, with the full record of attempts) than swap quietly. Within the tier we prefer the higher-quality route, so a failure can land you on a slightly lower-ranked one; every response carries `x-agiplan-model-actual`, so which model answered is never hidden.

No double billing

Failed attempts cost nothing

Only the attempt that actually returns is metered. Whatever we tried in between never reaches your bill.

Billed as served

Whichever line serves you is the one you pay for

After a failover you are billed at the price of the line that actually served the request, and the bill names that model. The attempts that failed never reach your bill — you shouldn't pay for our outage.

Every routing decision can be replayed

Automatic routing naturally invites the suspicion that you're being quietly served something cheap. The only answer is to show you every decision. Each request keeps a trace, retained 30 days, exportable.

EXAMPLEtrace tr_01J8FQ3M7XTotal 8.42s · Billed 4.970 CU
  1. 0msauthap_live_…f3c2 · 5× Pro
  2. 3msquotareserved 11.742 CU (14,200 in + max_tokens 8,192, worst case), all three windows clear
  3. 11msprismdifficulty 0.82 · tools=3 tokens_in=14,200 code_blocks=2
  4. 12msrouterule #1 matched → frontier · model-a
  5. 14msrespondmodel-a rate limited
  6. 210msfailoversame-tier backup model-b · billed at model-b's rate
  7. 224msstreamfirst token 890ms
  8. 8.42scommitin 14,200 / out 2,130 → 4.970 CU billed, remaining 6.772 CU of the reservation released

Traces hold metadata only, never request bodies — the same architecture as the zero-retention promise in section 07. When something breaks, send us the trace_id and support can pull it up instantly.

04Honest math

How quota is counted, out in the open

Price matters less than being able to check it. We use one unit, the CU: 1 CU = 1,000 output tokens on the reference model

Every model's input and output coefficients are published, and every call is inspectable, exportable and reconcilable. For a concrete sense of scale, here's what each plan is worth:

PlanMonthly quotaRoughly equivalent to
Lite30,000 CUAn entry-level monthly plan
Pro96,000 CUA standard subscription tier
5× Pro480,000 CU5× a standard subscription
Max1,920,000 CU20× a standard subscription
Three windows

5 hours · week · billing month

All three apply at once and are live in the console. The billing month is anchored to your signup date, not the calendar — subscribing on the 29th doesn't leave you with two days of quota.

Line by line

Every single call is inspectable

Timestamp, model, input and output tokens, CU charged, latency, trace ID. Exportable as CSV by project and date range.

No silent changes

Increases announced 30 days ahead

Terms for a subscription already sold never change mid-period. Plan versions are immutable; a price change applies only to new subscriptions.

Unused quota

Refunded pro rata

Cancel any time and get the unused portion of the period back. We don't make money on quota you never spend.

Live status

Whether each model is up right now

The status page gives you, per model: whether it is up right now, uptime over the last 30 days, whether it is available in your region, and every raw probe record (timestamp, outcome, latency, failure kind — exportable by month). Driven by real probes every 15 minutes. Multipliers live on /models. Latency is internal only for now; we do not publish it.

Quota alerts

We tell you before you run out

At 80% of your weekly or monthly quota (the threshold is fixed for now, not configurable) you get a heads-up plus a recommendation: which tier to switch to and how much longer it buys you, or what upgrading would cost. The 5-hour window sends no email — it resets several times a day, so mail would be noise; that figure is in every response header and on the live ring in your console.

05Pricing

We don't compete on price. That effort goes somewhere else: staying up, adding up, and owning it when things break.

Lite
¥49/mo

Enough to try it properly

  • 8,000 CU / week
  • 30,000 CU / month
  • 5-hour window 600 CU
  • Choose from every model
  • Prism auto routing
Get started
Pro
¥119/mo

≈ a standard subscription

  • 24,000 CU / week
  • 96,000 CU / month
  • 5-hour window 1,800 CU
  • Choose from every model
  • Email support, answered by a person
Get started
5× ProMost popular
¥569/mo

≈ 5× Pro

  • 120,000 CU / week
  • 480,000 CU / month
  • 5-hour window 9,000 CU
  • Custom Prism models
  • Priority routing (planned)
Get started
Max
¥1,368/mo

≈ 20× Pro

  • 480,000 CU / week
  • 1,920,000 CU / month
  • 5-hour window 30,000 CU
  • 8 concurrent requests
  • Dedicated routing (planned)
Get started
Annual billing is 20% off — two months freeVAT invoices available, bank transfer supportedWithin 7 dayswith under 10% used, full refundNeed more than your plan? Move to a cheaper tier, or step up a plan — see billing termsWorking with others? See team plans
06For teams

One contract, your whole team's usage and cost under control.

For teams the hard part usually isn't access — it's reconciling the bill at month end. All of this runs on our own gateway: quota, routing, metering and audit are one codebase, not middleware bolted together.

Central purchasing and billing

Buy by seat or by quota pack, invoiced once, paid by transfer. Annual master agreements and prepayment supported.

Department and project quotas

Split the pool across departments, projects and individuals, each with its own weekly and monthly ceilings and alert thresholds. One team overrunning never affects another.

Cost attribution

Consumption broken out by project, member and model, exportable to your own BI. Who spent what, at a glance.

SSO and member governancePlanned

Planned: OIDC and SAML plus the major workplace identity providers, joiner-leaver sync. Key scoping and forced rotation work today; single sign-on is not built yet.

A shared Prism policy

Define one routing policy for the team and push it to everyone. Costs stay predictable and nobody has to guess which model to use.

Private deployment and compliancePlanned

DPAs, security questionnaires and audit support are available today. Gateway nodes inside your own VPC are not available yet — talk to us first rather than planning around them.

07Your data

Your code and prompts are not stored by default

You're about to send us the most sensitive material you have — source code, internal documents, client data. This section is wordier than a marketing page should be. Please actually read it.

Zero retention

Request and response bodies never hit disk

The gateway records only what metering requires: model, token counts, latency, status code, trace ID. Bodies are forwarded in memory and released the moment the stream ends — no logs, no database, no object storage.

Diagnostic logs

We haven't built this switch

Seeing the body would obviously help when debugging, and we have not built it — bodies are never written to disk, with no exception. If we ever do build it, it will be off by default, enabled only by you, bounded by a stated retention period, and announced in advance. Until then there is no path at all by which we could see your bodies.

Never trained on

No training, no resale, no human reading

Your content never enters any training or fine-tuning pipeline, is never handed to third parties, and staff have no access. All three are written into the terms of service and legally binding.

In transit

Encrypted end to end

TLS 1.3 externally, mutual TLS between internal services, and encryption enforced on upstream connections too.

Keys

Hashed — we can't read them either

API keys are stored hashed and shown in plaintext exactly once, at creation — we can't read them back either, so a lost key is rotated rather than recovered.

Access control

Least privilege, fully audited

Every back-office action touching an account is recorded: who, when, what changed, from which IP. Audit logs are tamper-evident and can be produced on request.

Your rights

Export or delete at any time

Usage data is exportable whenever you want. On closure, your email, credentials, keys and technical logs are destroyed within 30 days; the transaction and usage ledgers cannot be deleted and are instead unlinked from you — those three record types are append-only, and a ledger that can be edited proves nothing in a billing dispute. The privacy policy spells this out. (Account closure is not built yet; see the terms below.)

If something happens

You hear from us within 72 hours

If a security incident could affect your data, we notify affected users within 72 hours and publish a post-mortem.

08Aligned incentives

How we make money, stated plainly

The most reliable way to predict whether a company will treat you badly is to understand how it earns. So here it is.

We earn the spread on purchasing scale and routing efficiency.

We buy at volumes far beyond any individual user and turn that scale, plus Prism's routing efficiency, into your price. This business only works if you stay: acquiring a customer costs far more than a month of gross margin, so burning you once is a net loss.

We don't earn on quota you never spend — cancellations are refunded pro rata. We don't earn by giving you less than you asked for — the model you pick is the model that runs. We don't earn on your data — no training, no resale.

Long term the profit comes from teams and enterprises, and enterprise buyers only work with vendors that have a clean reputation. So there's no goodwill involved here. We just did the math.

Eight things we will never do

This list becomes a standalone Service Commitments page and takes effect alongside the terms. Break any one of them and you can demand a full refund.

  1. 01

    Never silently substitute the model you chose

    No reduced parameters, no truncated context, no quiet downgrade at peak hours. The response header carries the model that actually ran, so you can check for yourself.

  2. 02

    Never train on, resell, or read your content

    Your prompts and outputs enter no training pipeline, go to no third party, and staff have no access.

  3. 03

    Never reduce quota terms mid-subscription

    Any change to multipliers or quotas applies to new periods only; increases are announced 30 days in advance.

  4. 04

    Never use countdowns, scarcity, or referral bounties

    The price is the price. A countdown exists to stop you thinking it over, and we don't need that conversion.

  5. 05

    Never claim official authorisation from any provider

    We are an independent aggregation gateway and will not say otherwise. When another platform does, ask them for it in writing.

  6. 06

    Never pretend an outage didn't happen

    Any incident affecting more than 1% of users gets a signed post-mortem within 48 hours: what happened, root cause, what we changed.

  7. 07

    Never put obstacles in front of a refund

    Refund rules are fixed in the terms and self-service in the console. No support ticket, no reason required, no retention script.

  8. 08

    Never disappear without notice

    If we ever shut down, you get 60 days' notice, a full refund of unused quota, and help migrating elsewhere.

09What we oppose

Your request bodies have a market price

This rarely gets said out loud. Every prompt and every file a user sends can be recorded in full — distilled into smaller models, packaged and sold as training data, or quietly traded between platforms. No notification is sent, and the terms usually contain no explicit prohibition.

The internal documents, client data and unreleased code you pasted in may already be in somebody's training set.

We're against this, and not only in words.

  1. 01

    Zero retention is architecture, not policy

    There is no code path in the gateway that writes a request body anywhere. Not "we promise not to look" — there is nowhere to put it. This is auditable.

  2. 02

    We buy no data of unclear origin

    We purchase request corpora from no platform and no data broker, at any price, under any label.

  3. 03

    We refuse data-for-price deals

    When a provider offers a lower rate in exchange for data sharing, we decline — even though it means higher costs and thinner margins for us.

  4. 04

    We're willing to pay for it

    Every line above reduces what we can earn. That's the price we chose to pay, and the reason we'd like your money instead of somebody else's.

  5. 05

    We accept verification

    All of this goes into the terms of service with legal force. We accept third-party security audits, and you're welcome to inspect the traffic yourself.

We can't change the whole industry. But whatever passes through our pipe, we can guarantee it goes nowhere else.

10Who we are

This is not an anonymous site.

Most services like this are anonymous: no legal entity, no names, a domain that can change overnight. Our names, engineering notes and post-mortems are all here for you to check.

Name TBDGateway & meteringQuota reservation, Prism routing, billing reconciliation
Name TBDProduct & designConsole, documentation, everything user-facing
Name TBDSupply & operationsMulti-provider purchasing, capacity planning, availability
Name TBDSupport & complianceSupport email, invoices, contracts, security questionnaires
Engineering log — weekly, every entry signedStatus page — full 90-day historyPost-mortems — all public
11What worries you

You're probably wondering.

I've never used anything like this. Can I set it up myself?

Yes. Sign up, create a key in the console, and run one line that writes it into your local config. Then open the coding tool you already use — no config files to edit by hand. If you get stuck, open a ticket in the console and a person answers. We email you when a ticket gets a reply (the email carries no reply text — the content stays on the site).

What if you shut down in six months?

That's the reasonable worry, and a promise alone is worth nothing. What we can do is make leaving expensive for us: legal entity and registration published in the footer, invoices, contracts, named team. Service Commitment 8 states that if we stop, you get 60 days' notice, a full refund of unused quota and migration help. You can also start on ¥49 Lite for a month — there's no annual lock-in.

Could my quota be quietly reduced?

No. Plan versions are immutable: the quota terms and multipliers at the moment you subscribed are fixed for the whole period, and any change applies to new periods only with 7 days' notice. That's a database-level constraint, not just a promise — what's stored is the version you bought.

With Prism routing, am I quietly getting cheap models?

Yes — but only when the task doesn't need a stronger one. The logic and the outcome are both visible to you: every request's trace carries the difficulty score, the rule that matched and the model that actually ran, retained 30 days and exportable. Borderline difficulty always rounds up, and requests declaring tool use or long context exclude models that can't serve them. You can also skip Prism entirely and name the model yourself.

What happens when a model goes down? Do I pay more?

Failover is on by default: Prism moves the request to a backup model in the same capability tier, so your call doesn't fail. Failed attempts aren't billed — only the one that returns. If the only backup costs more, you're still billed at the tier you requested and we absorb the difference. You can turn failover off (set failover.enabled to false in a mix); you can't currently choose the backups or their order — Prism picks within the same tier.

What if I run out of quota early?

Two options today: switch to a cheaper capability tier and keep going; or upgrade, with the difference prorated over the remaining days and your billing date unchanged. There is a third route: pay-as-you-go, drawing from your balance at ¥1 = 500 CU with the same per-model rates as your plan. Off by default; enabling it requires a monthly ceiling, and charging stops there.

Can anyone read my code and prompts?

Not stored by default. The gateway records metering metadata only; bodies are forwarded in memory and released when the stream ends. No training, no resale, no staff access — all three are written into the terms with legal force.

Is every model available everywhere?

No. Some models have not completed generative-AI service filing in mainland China and are therefore not offered to users there; the console labels regional availability per model. Please state your location accurately — circumventing regional restrictions is the user's own responsibility, as set out in section 4 of the terms.

Why you instead of buying from the provider directly?

Three reasons: CNY billing and invoices, so no credit card or FX; one quota covering models from several providers including domestic ones, switchable without code changes; and Prism, which stretches the same money further — how much further depends on how much of your work genuinely needs the frontier tier: below roughly 30% frontier share the same quota does about twice the work; if 90% of your calls must run on frontier, the difference is small. The comparator on /prism lets you work it out yourself. If you use exactly one model from one provider and already have a working payment method, going direct is simpler. We don't think we need to convince everyone.

Are you officially authorised?

No, and we won't claim to be. We are an independent multi-provider aggregation gateway. When a comparable platform claims official authorisation, ask them to show it in writing.

Compute should work like utilities:metered transparently, drawn on demand,and switchable to another supplier at will.

A credit card, an exchange rate, a regional restriction — none of these should be the reason someone can't use the best AI available. That's why we started.

Then we found a second reason. Good models are becoming accessible, but the price is often that you hand over your data without knowing it. That price shouldn't exist, and we'll spend what we have to hold it back.

What we build is the pipe in the middle. We want it sturdy, we want the accounting honest, and we want you free to leave whenever you like. While you're using it, it's only a pipe — it doesn't keep a single word.

Try it for a month, then decide whether to trust us

No annual lock-in, cancel whenever. Full refund within 7 days if you've used under 10% — self-service in the console, no ticket required.