← Back to home

Engineering log

Tripping the breaker is easy; where you go next is not

In the 19 July incident the circuit breaker worked exactly as designed. What failed was the reserve capacity sitting behind it.

The "19 July incident" this piece describes is a placeholder — it never happened. It is the same entry that carries the SAMPLE badge on the incidents page. It is written this way to show how we intend to write engineering logs: timelines to the minute, causes stated plainly, changes signed. A real incident replaces it, and this notice goes away.

Last updated
2026-07-24
Author
Supply & operations

"Timeout" is the wrong breaker threshold

Our original condition was "the request timed out". It sounds reasonable and it is far too late: by the time a 20-second timeout fires the user has waited 20 seconds, and a queue of similar requests has already built up.

What actually runs today is "two consecutive failures", backing off from 30 seconds up to 15 minutes. Only 5xx, connection failures and timeouts count — 429 does not: being rate-limited means the route is alive, and tripping it just pushes load onto everyone else, which is the cascade we are trying to avoid. What we want is "first-token latency above 3× the rolling median for 60 seconds": normal latency differs several-fold between models, so a fixed threshold is either hair-trigger or useless. It is not built yet — p50/p95 latency currently only feeds route scoring, not the breaker. This paragraph used to end with "the change moved the trip roughly 9 minutes earlier". We never measured that. The Chinese version of this article says, in the same place, that writing a number you have not measured takes the whole piece down with it — and then the English version wrote one anyway.

After it trips, is there enough capacity

The moment the breaker moves all of A's traffic onto B, the question stops being "is A healthy" and becomes "can B carry it". If B's reserve concurrency is provisioned with only a little headroom over normal peak, nothing fails outright and everyone queues — which is harder to diagnose than an error, because nothing on the dashboard turns red.

"There is a backup" and "the backup can carry it" are different claims. What is needed is a capacity rule: alert when any tier's reserve concurrency falls below some multiple of current peak, before it is ever needed. That is not built yet — and we are not going to invent the multiple here either; it has to be measured from a real peak distribution. The previous section said writing an unmeasured number collapses the credibility of the whole piece; the same applies here.

Why we don't drop across tiers by default

Cross-tier fallback is tempting: nothing left in this tier, so surely a step down beats failing. We chose not to do it by default.

Because the user has no way to verify the downgrade was necessary. Allow silent downgrades once and the suspicion "am I quietly being served something cheap" never washes out — which is exactly the weakest point of the whole Prism proposition. So when a tier has nothing available the default is a 503 with the full record of attempts. Users who want cross-tier fallback can turn it on explicitly.

The cost is more 503s than we'd otherwise have. We accept it, because verifiability is worth more than uptime.

How failover is billed

  • Failed attempts aren't billed. Only the attempt that returns is metered; whatever we tried in between never reaches the user's bill. But once a stream has started and the upstream dies, what was already produced is billed — it runs the same teardown as a user-initiated abort (`settle("error")` and `settle("aborted")` both commit real usage).
  • Switching to something cheaper bills the cheaper rate. The user comes out ahead and we don't claw it back.
  • Switching to something pricier also bills what was used. Every line has its own price; whichever line serves the request is the one billed, and the bill names that model.

That last one is a real cost, and it lands precisely when we're already dealing with an outage. But charging a user extra because we broke something would destroy trust in a single incident — we did the arithmetic, and we can afford it.