Post-mortems

Everything that went wrong, none of it buried

For any incident affecting more than 1% of users we publish a signed post-mortem within 48 hours. Impact is stated as a specific percentage — "some users" says nothing, and you can tell when someone is hedging.

The three entries below are placeholders — none of them happened. They are written this way to show how we intend to write post-mortems: impact as a specific percentage, timelines to the minute, cause and fixes signed. Real incidents replace them and drop the sample badge. We are not going to leave this page empty — a post-mortem page that never has anything on it only means nobody is writing them.

SAMPLE2026-07-19 14:22 CSTDuration 47 minImpact 3.1%Owner Gateway & metering

Elevated first-token latency on the frontier tier, some requests failed over

Frontier

One upstream provider's latency rose from a 1.2s median to over 9s, past half our first-token timeout. Failover moved requests to a same-tier backup as designed, so nothing failed outright, but affected users noticed a clear slowdown.

Timeline

  1. 14:22Probes detected frontier-a first-token latency at 3× baseline; weight automatically reduced.
  2. 14:31Latency kept climbing; the circuit breaker tripped and all traffic moved to frontier-b.
  3. 14:38frontier-b approached its concurrency ceiling and queuing appeared. We intervened and temporarily lifted the cross-provider cap.
  4. 15:09Upstream recovered; traffic was shifted back gradually and confirmed stable after 20 minutes of observation.

Root cause

We were too optimistic about single-provider capacity. frontier-b's reserve concurrency was provisioned at only 1.3× normal peak — not enough to carry the tier alone when the other provider went down. The breaker worked correctly; the reserve behind it didn't.

What we changed

  • Reserve concurrency on the frontier tier raised from 1.3× to 2.5× normal peak, and folded into the capacity-forecast alerting.
  • Breaker threshold refined from "timeout" to "latency above 3× baseline for 60 seconds", firing roughly 9 minutes earlier.
  • Every request that failed over during the window was billed at the originally requested tier, with the difference absorbed by us and nothing to claim.
SAMPLE2026-05-30 03:10 CSTDuration 132 minImpact 1.4%Owner Gateway & metering

Usage records lagged, console figures roughly two hours behind

The async job that persists usage events backed up, so the console showed less quota used than reality. Quota enforcement was unaffected — the Redis counters were correct and rate limiting followed real usage. Only the displayed detail lagged.

Timeline

  1. 03:10Persistence queue backlog alert fired.
  2. 03:52Traced to a query missing an index that was slowing batch writes.
  3. 05:22Index added, backlog cleared, reconciliation confirmed no data loss.

Root cause

A newly added per-project filter introduced a query without a covering index. At end-of-month data volumes it degraded to a full scan and contended for locks with the persistence job.

What we changed

  • Index added, plus a slow-query regression check in CI.
  • When the backlog exceeds 5 minutes the console now shows a "data delayed" banner, instead of quietly showing a wrong number.
SAMPLE2026-04-02 11:05 CSTDuration 26 minImpact 0.6%Owner Support & compliance

Region filter misfired and wrongly blocked some users in mainland China

Domestic flagship

A region-rule update wrongly marked a batch of models as unavailable in mainland China, and affected users saw them greyed out. No incorrect billing resulted, and there was no inverse failure — nothing that should have been blocked was let through.

Timeline

  1. 11:05Rule published.
  2. 11:18A support ticket reported models unselectable; on-call confirmed a misclassification.
  3. 11:31Rule rolled back, service normal.

Root cause

The filing-status field defaulted to "not filed" rather than "unchanged", so any model not explicitly named in the update was marked unavailable too.

What we changed

  • The field is now mandatory; an update that leaves it unset is rejected at publish time.
  • Region-rule publishing now shows a diff preview listing every model whose availability the change would alter.

Once live, this page is driven by real events.