inferrail

0021. Atomic budget reservations, streamed-call accounting, and an explicit request-field policy

Status

Accepted (pending review)

Context

ADR-0015’s block-mode budgets check spent_so_far + estimate <= limit and then call the provider. Spend is only recorded when the receipt is written, after the provider call. That is a check-then-act race: when several requests share a budget scope (parallel tool calls, sub-agents, concurrent runs of one job under one work_id), each can pass the check before any of them records spend, and together they spend past the limit. The same gap applies across gateway processes sharing one budgets database.

Three related gaps make a per-run dollar ceiling for an agent run less trustworthy than it should be:

  1. Concurrency — the race above.
  2. Streamed calls with no reported usage — an OpenAI-shaped stream only carries usage if stream_options.include_usage is set. Without it, the receipt’s cost is null (honest), but that spend then counts as $0 against the budget, so a block budget quietly stops limiting anything.
  3. Request fields — ChatCompletionRequest rejects every field it doesn’t model. Agent frameworks send several provider-valid fields (response_format, max_completion_tokens, store, metadata, reasoning_effort, the developer role, text content-part arrays, refusal on assistant history messages), so requests fail before a budget is ever involved.

Decision

Reservations (admission is atomic)

For every budget that matches a request (global, project, work_id):

available = limit − committed spend − outstanding reservations

Settlement (after the provider call)

A reservation is settled per provider attempt. Each retry re-admits and reserves again, and can itself be refused.

Outcome of the attempt Settlement
Usage reported and priced (success, or a partial stream whose final usage arrived) Receipt written first, then the reservation is released. There is a brief window in which both are counted, which errs toward refusing and never toward overspending.
Provider answered with an HTTP error before any output Released (the provider rejected the request)
Timeout or transport failure (the request may have reached the provider) Held
Stream ended, was cancelled, or failed without reported usage Held
Usage reported but the cost can’t be computed exactly Held

A held reservation keeps counting against the budget at its reserved amount. It is never shown as the request’s cost: the receipt’s estimated_cost_usd stays null, and the receipt carries a separate budget_held_usd attribute with the amount held. An estimate is never presented as an exact cost.

Unknown pricing

With a block-mode budget in scope, a request for a model with no verified price is refused before the provider call (INFERRAIL_E012, HTTP 402). There is no dollar amount to reserve, and letting the request through would make the ceiling meaningless. Warn-mode budgets keep logging only. The operator can add a pricing override for the model.

Crash and restart

Reservations are persisted, and they are never released automatically at startup, because another gateway process may still own them. After a crash, an active reservation keeps counting (fail-closed) until its budget window rolls over. For a per_work budget that means the rest of that work item’s life.

A cancelled attempt (for example the request task is cancelled while the provider call is in flight) is held, not left active. One known gap remains: ADR-0006 describes a narrow race in which the server discards a streaming response before it is ever iterated. In that case no settlement code runs, so the reservation stays active and keeps counting (fail-closed).

Streaming usage

For a verified OpenAI provider the gateway already requested usage on streams unless the caller set stream_options. It now requests it whenever the caller didn’t set include_usage itself, so a caller who passes other stream_options keys still gets a priced receipt. A caller’s explicit include_usage: false wins, and that stream is then held as above.

Request-field policy (/v1/chat/completions)

Top-level fields fall into three explicit lists. Anything not listed is still rejected with INFERRAIL_E006: this is not blind pass-through.

Messages accept the developer role, name, refusal on assistant messages, and content either as a string or as an array of {"type": "text"} parts. Other part types and unknown message keys are now rejected; before this change they were silently dropped. A non-streaming response now also carries the model’s refusal, so structured-output clients see it.

Consequences