Quick start
After signing up, run one line. The switcher installs itself, walks you through browser authorisation, and points your local tools at us. Then open the coding tool you already use and carry on.
curl -fsSL https://agiplan.dev/setup | sh agiplan status # what's left in each of the three windows agiplan models # which models are available eval "$(agiplan env)" # point your local tools at us
You can skip the script entirely. Two environment variables do the same job — it's only less convenient when you juggle several configurations.
export API_BASE_URL="https://agiplan.dev/v1" export API_KEY="ap_live_••••••••••••••••"
What the script actually does
`curl | sh` asks you to trust us, so here is every change it makes. The script is served as plain text — open https://agiplan.dev/setup in a browser and read the whole thing. What you read is what you are about to run.
- Writes the API base URL and your key to `~/.config/agiplan/env` — directory 700, file 600, umask tightened before the write so the key is never briefly world-readable. It doesn't touch system directories and doesn't ask for sudo.
- The same file also exports `OPENAI_BASE_URL` / `OPENAI_API_KEY`. Most tools read those names, so once it is sourced they will route through us — that is intended, but you should know it.
- It does not modify your `.bashrc` / `.zshrc`. Whether it loads at login is your call, and that is your file.
- It installs no binary. The CLI isn't published yet and the script doesn't pretend otherwise — an install script that fakes success is far more dangerous than one that states where things stand.
- It does not detect your installed tools, does not modify any other config file, and uploads nothing about your machine.
- If the key doesn't start with `ap_live_`, it exits and writes no file at all.
The script does not touch your shell config — it writes exactly one file, `~/.config/agiplan/env`, so there is no backup to restore. `agiplan uninstall` deletes that one file, and by default only lists what it would delete; pass `--yes` to actually do it.
Authentication
Both protocols use a bearer token. Keys are created in the console and the plaintext is shown exactly once — we store only an argon2id hash, so a key can't be recovered, only rotated.
Authorization: Bearer ap_live_xxxxxxxxxxxx
Keys can be scoped, bound to a custom model, IP-restricted and given an expiry. Rotation is deliberately easier than revocation: the correct action should be the least effort. Old keys can keep a grace period rather than being cut off instantly.
Which protocols
The same base URL serves compatible endpoints for three mainstream protocols. Which one you use depends on your SDK — you don't have to tell us.
| Endpoint | Protocol | Notes |
|---|---|---|
| /v1/chat/completions | OpenAI-style | stream, tools, response_format |
| /v1/messages | Anthropic-style | stream, tools, system |
| /v1/responses | OpenAI Responses | stream, tools; `previous_response_id` avoids resending context |
| /v1/models | Any | Returns the models and custom models available to you |
| /v1/quota | Any | Read the three windows without making a call; needs usage:read |
Any tool that lets you change its endpoint just works — no plugin, no code changes. Differences between the request shapes are absorbed inside the gateway; write whichever you already know.
`/v1/responses` is the only stateful one. When you use `previous_response_id`, that conversation's messages are stored on our side for at most 5 hours and then deleted by a scheduled job; not using the parameter produces no such storage. We hold that state ourselves rather than at an upstream so failover stays invisible to you — those 5 hours are the price, and the privacy policy says the same thing.
A few endpoints are deliberately not in the table: `/v1/device/code`, `/v1/device/token` and `/v1/keys/current` are the CLI's transport (used by `agiplan login` and `agiplan use`), not something you call directly. `/v1/traces/{id}` is covered under tracing below.
Images work as inline `data:` URLs (base64) only — remote `http(s)` image links are not built yet. Models tagged "vision" really can see images, but you have to carry the bytes in the request body. Given a URL, the gateway would have to fetch it and re-encode it while blocking internal addresses (SSRF), and we have not finished that — so for now it is downgraded to a text placeholder. When that happens the response carries `x-agiplan-dropped-params: content[].image_url`: we don't want you believing the model looked at an image it never saw. Audio and file parts behave the same way — currently dropped, and listed in the same header.
Name a model, or hand it to Prism
One `model` field, two ways to use it: name a model and it's pinned; put `auto` or your own custom model name there and Prism routes by task difficulty.
"model": "frontier-a" // pinned, Prism stays out of it "model": "auto" // handed to Prism "model": "my-coding-model" // your own custom model
To override the difficulty judgement on a single request, add a header: `X-AGIPlan-Difficulty: 0.9`. Every model and coefficient is listed at /models.
Streaming
Both protocols stream SSE in their own native format, behaving exactly as you'd expect. The one thing to know: billing settles when the stream ends.
On arrival we reserve quota for the worst case — the output side from the `max_tokens` you declared, the input side from the tokens already countable in the request. A normal finish, a client abort and an upstream error all run the same teardown, reconciling against actual usage and returning the difference. Disconnecting midway never bills you the worst case.
Reading the quota response
Every response carries the remaining percentage and reset time for all three windows. Percentages rather than counts is deliberate: enough to decide whether to back off, without writing your plan's absolute allowance into every response header. To reconcile exactly what you spent, use the usage detail in the console or the monthly statement — those are absolute. On rate limiting you get a 429 that says which window is full and when it recovers — the 429 carries the same three headers, which is exactly when you need them. An invalid key (401) carries none of them: at that point we don't yet know whose quota to report.
X-AGIPlan-Quota-Session: remaining=68%; resets=2026-08-15T09:00:00.000Z X-AGIPlan-Quota-Week: remaining=34%; resets=2026-08-18T00:00:00.000Z X-AGIPlan-Quota-Month: remaining=81%; resets=2026-09-05T00:00:00.000Z Retry-After: 1740 // seconds, only on 429
Errors and retries
| Status | Error code | Meaning | What to do |
|---|---|---|---|
| 401 | invalid_api_key | Key invalid or revoked | Use a different key; do not retry |
| 404 | model_not_found | Images endpoint: no such image model | See the entries with kind image in /v1/models |
| 503 | size_unavailable | Images endpoint: no route supports that size right now | Try another size or retry later. Failed calls are not billed |
| 401 | key_rotation_required | Key is older than the org's rotation policy | Rotate a new one in the console |
| 400 | capability_unavailable | No connected model supports a capability the request declares (vision, tools…) | **Retrying will not help.** Drop the requirement, or wait for an upstream that supports it |
| 400 | context_too_small | The model you named by its raw ID has a smaller context window than this request's `min_context_k` | **Retrying will not change the result.** Shorten the context, or pick a model with a larger window — /models lists the window for every model |
| 400 | upstream_rejected_request | The upstream rejected the request itself (parameters or content) | **Another model gives the same answer.** Check the request; if you believe it is valid, send us the trace id |
| 400 | request_exceeds_mix_cap | Worst-case estimate exceeds the `max_cu_per_request` set in your custom model | **This is your own cap.** Lower max_tokens, shorten the context, or raise the value |
| 403 | insufficient_scope | Key lacks the required scope | Use a key that has it |
| 403 | no_subscription | No active subscription on this account | Subscribe or renew in the console |
| 429 | quota_exceeded | A window's quota is exhausted | Wait per Retry-After, or move to a cheaper capability tier |
| 429 | concurrency_exceeded | In-flight requests hit your plan's ceiling | Let one finish. **Upgrading fixes this; topping up does not** — it is not a quota problem |
| 451 | region_not_available | No model in that tier is offered in your region | **Retrying will not change the result.** Pick another tier, or see the region column at /models |
| 503 | no_supply_configured | No upstream is connected for that capability tier | **Retrying will not change the result.** Pick another tier, or use auto — /models shows how many lines each tier has right now |
| 503 | policy_unavailable | The custom model you called no longer compiles, usually because a model it pins was withdrawn | The error message names the pool and the model. In the console under Custom models, point that pool at another model or a tier; it takes effect as soon as you save |
| 503 | model_unavailable | The model you named by its raw ID has no working route right now | **We will not substitute another model** — you pinned a raw name, and swapping it silently would be lying to you. Retry later, or use a capability tier name (e.g. `domestic-a`) so the gateway can switch between equivalent lines |
| 503 | all_upstreams_failed | Every model in the tier was tried and none worked | The body carries the full record of attempts. Failed attempts are not billed |
| mid-stream | upstream_interrupted | The upstream dropped after output had started. The status is already 200, so this arrives as an error event inside the stream, **with no normal end marker after it** | **Output already sent is billed by actual tokens.** Switching models mid-answer would make it inconsistent, so the gateway does not fail over here; resend the request |
You don't need to retry 5xx yourself. Failover already ran inside the gateway — a 503 means every model in the tier was tried. Retrying from the client only adds load upstream.
This table lists the codes the gateway really returns. Running out of quota is still a 429; with pay-as-you-go on, the request draws from your balance instead of being refused. If the balance can't cover it either, it is still a 429 and the message says which gate failed (not enabled, ceiling reached, insufficient balance). A test pins this table to the codes in the code.
Branch on `code`, not on `message`. `code` is the stable contract and changes are announced; `message` is a sentence for a human and its wording will change. One thing worth stating plainly: `message` is currently Chinese only — the site, the docs, the console and our notification emails are all bilingual, this one string is not yet. `code` and the HTTP status are unaffected, so code that branches on them is not affected by the language.
Trace IDs
Successful responses carry a trace ID and the model that actually ran. If failover happened, that's marked too. Refused responses carry no trace ID — that call never became a usage record, and handing you an ID that resolves to nothing is worse than none. On a 503 the chain is in the response body as `attempts`; send that along when you report the problem.
X-AGIPlan-Trace: tr_01J8FQ3M7X X-AGIPlan-Model-Actual: frontier-b X-AGIPlan-Failover: true
Fetch the full trace with your own key: `GET /v1/traces/{id}`, structured JSON, retained 30 days. Send the trace ID when reporting a problem and support pulls it up directly — you don't have to reconstruct what happened.
Command line
First, the thing that matters more: the CLI has no distribution channel yet. It is not on npm, there is no download, and the setup script says outright that it installs no binary. So the table below records implementation status — which commands are written, which one we decided against — not whether you can run them today. Right now `agiplan` gives you command not found, and that is not your configuration.
"Not built", "broken" and "built but you cannot get it" are three different things, and conflating them wastes the most time. Commands that are not built say so plainly when run; the fact that the tool itself is unreleased can only be written here. Once we decide how to ship it (npm, a single binary, or bundled with the setup script) this paragraph becomes install instructions.
| Command | What it does | Status |
|---|---|---|
| agiplan login | device-code auth: confirm in the browser, key delivered to the terminal | Available |
| agiplan status | remaining percentage in each of the three windows | Available |
| agiplan models | available models and capability groups | Available |
| agiplan use <name> | bind a custom model to the current key; --none unbinds | Available |
| agiplan env | print evaluable environment variables | Available |
| agiplan keys rotate <config> | rotate a key | Declined, see below |
| agiplan uninstall | delete the config file setup wrote (--yes to actually delete) | Available |
We decided not to build `keys rotate`, for security rather than effort. A key that can issue its own successor outlives revocation: after a leak the attacker rotates once, you revoke the old key in the console, and the new one is still alive — and you don't know it exists. Rotation only happens through a browser-confirmed path: rotate in the console, or run `agiplan login` again.
A per-project `.agiplan` file is supported. It holds only the profile name, never a key, so committing it is safe.
Migrating from somewhere else
If you already use an official SDK or another aggregator, only two things change: the endpoint and the key. Where model names differ, a custom model name gives you a mapping layer and your application code stays untouched.
One behaviour differs from the official API and deserves its own paragraph: if you omit `max_tokens`, we substitute 4096 and send that upstream. Officially, omitting it means "use the model maximum", so the same code pointed here will see long answers truncated at 4096 with `finish_reason: "length"` — looking as though the model stopped on its own. We do it because that number also determines how much quota we reserve: with no ceiling there is no way to check your balance before the call. When we substitute it the response carries `x-agiplan-max-tokens-default: 4096`; send your own value and yours is used.
One more thing that changes your bill, written here rather than left for you to discover: when a request needs a capability the tier you picked cannot provide, we lift it to a tier that can — and bill the tier it landed on. The common case is images: your custom model selects the fast tier, no model there has vision, so the call lands on the domestic flagship tier and the output coefficient goes from 0.06 to 0.18 — three times the CU for the same tokens. We lift rather than fail because failing serves you less; and we only ever lift, never drop, so a downgrade never happens silently. Every response carries `x-agiplan-billed-group` naming the tier actually billed, matching what appears in your usage detail. To avoid being lifted, don't send what that tier can't handle, or pick a tier that supports it outright.
- Create a custom model named after the model name you use today, routing to the matching capability tier.
- Change the base URL and key, then run your tests.
- Check the coefficients at /models and reconcile your first day against the usage detail in the console.
Don't skip step three. Reconciling on day one is far less work than discovering a discrepancy later and going digging — and we'd like you in the habit.
The docs are still being filled in. If something's missing or unclear, email support@agiplan.dev — we'll fix it and note it in the changelog.