← Back to home

Engineering log

When exactly should a streaming request be charged

You don't know the cost when the request starts, and by the time it ends the money is already spent. This is the easiest thing to get wrong in a metering system. We solve it in two phases.

Last updated
2026-08-06
Author
Gateway & metering

The problem: both obvious approaches are wrong

Metering purely after the fact is the intuitive one: the stream ends, count the tokens, deduct. The trouble is it can't prevent overrun. A user with 100 CU left can fire ten large requests at once; every one passes the check on entry and together they drive the balance negative.

Charging upfront against `max_tokens` goes to the opposite extreme: a user sets `max_tokens: 8000`, generates 200 tokens, and is billed for 8000. That pushes people to set max_tokens far too low, which truncates long answers — we'd have damaged output quality with billing logic.

The fix: reserve, then commit

The standard answer is two phases. Reserve the worst case when the request arrives, then commit against real usage when the stream ends, returning the difference.

1. auth      api_key → user → subscription → three window limits
2. reserve   atomic Redis Lua script (Postgres advisory lock when Redis is off):
               check(5h, week, month) && hold(estimate) && zadd(reservation, TTL=15min)
             estimate = max_tokens × output coefficient + tokens_in × input coefficient
3. forward   pick upstream → stream SSE through
4. persist   write usage_events (append-only) — **truth first**
5. commit    real usage → release the hold, add to the counters, settle the difference
6. reconcile a scheduled job recomputes counters from usage_events:
               record the drift first, then repair the cache

Step 2 has to be atomic. If checking three windows and decrementing takes several round trips, two concurrent requests can both pass the check. A Lua script runs single-threaded inside Redis, which solves it for free.

With no Redis configured this step uses a Postgres transaction-level advisory lock, and one user's concurrent requests queue on it. The verdict is identical; the cost is two extra aggregate scans per request. A set of tests compares the two paths case by case, so that configuring Redis can never quietly change the business rules.

The order of steps 4 and 5 is part of the correctness argument: truth is persisted first, counters follow. The other way round, a failed write leaves usage nobody can reconcile; in this order, a failed counter update at worst leaves the cache low, and the next warm-up reads the true total — which already includes that call — straight back out of usage_events.

The hard part is the edge cases

The happy path is twenty lines. What's hard is what happens when it doesn't. A few we've hit:

  • The client disconnects midway — the most common one. Teardown has to hang off `finally`, not the stream's normal completion callback, or the reservation stays held.
  • The gateway process dies — `finally` won't save you either. So reservations carry a 15-minute TTL: the worst case is a user's quota being briefly held, not permanently lost. That's the direction we're willing to fail in.
  • Upstream usage disagrees with our count — upstream wins, and the difference is recorded as its own metric. A persistent drift means our tokenizer assumption is wrong and someone needs to look.
  • Redis is down — fall back to the Postgres advisory-lock path and count the degradation. Not fail-open: letting requests through means anyone can exceed their quota for a while, and that excess lands in an append-only usage table we cannot claw back. Slow is observable and can be scaled; fail-open is a loss you discover afterwards.

Redis is fast; Postgres is the truth

The Redis counters exist for speed, but they drift: a crash, an early TTL expiry, a bug in the Lua script — any of them moves the count away from reality. So the truth lives in `usage_events` in Postgres: append-only, never updated.

A scheduled job recomputes the counters from the event table. When it finds a discrepancy it writes the drift to an append-only table first, and only then repairs the cache — the cache, never the books. The record matters because the gap itself is the diagnostic: it is the only thing that can later answer whether users were over-admitted during those minutes.

This is deliberately different from balance reconciliation, which only reports. The difference is not which table matters more, it is whether repairing destroys the evidence. The ledger is the truth, so correcting it erases the trail; the counters are a cache, and leaving them wrong keeps making wrong admission decisions — while the trail already sits in the drift table.

The same principle runs through the whole billing system. Balance works this way too: `balance_ledger` is the truth and the balance column is only its materialised snapshot. Any code path that changes a balance without writing to the ledger is one we won't write.