inferrail

0022. Per-run budgets without pre-registration: a declared or default ceiling per work_id

Status

Accepted (pending review)

Context

ADR-0021 made block-mode admission atomic. But protecting a single agent run still means creating a budget object for its work_id first (inferrail budget set --scope work_id --scope-value <id> …), and doing that again for every new run. Most runs get their id when they start, so this setup step is the main friction in “give this run a dollar ceiling”.

The desired flow:

mark the work (work_id)
  → declare a budget, or inherit a default
  → no separate setup
  → enforced at admission
  → a final per-run record

Decision

How a run gets a budget

A request that carries X-Inferrail-Attribute-Work-Id: <id> can be covered by a per_work, block-mode budget for that id in three ways:

  1. Stored: a budget already exists for work_id:<id>:per_work (created with the CLI or local API, or created earlier by 2 or 3 below).
  2. Declared: the request sends X-Inferrail-Budget-Usd: <amount>, and no budget exists yet for that id.
  3. Default: the operator set budgets.per_work_default_usd, no budget exists yet for that id, and the request doesn’t declare one.

With 2 or 3, the budget row is created inside the same admission transaction that reserves the request (ADR-0021’s BEGIN IMMEDIATE). So when concurrent first requests for a new id all declare the same budget, exactly one creates it, and every one of them is admitted against it. No request for an unseen id can slip past before the budget exists.

A request with no work_id isn’t affected. Neither is a work_id request when no budget is stored, declared, or defaulted. Global and project budgets still apply exactly as before.

Override hierarchy and immutability

Trust boundary and spoofing

Nested runs, retries, streaming, pricing, failures

Running in front of another gateway (coexistence)

A team can keep its current OpenAI-compatible gateway and add Inferrail in front of it just for per-run control. This was tested locally in front of LiteLLM and otari; other gateways haven’t been tested:

framework → Inferrail → existing gateway → provider

Two opt-in settings on an openai_compatible provider remove the friction found in testing:

Each gateway enforces only its own budgets, so nothing is counted twice.

A downstream budget refusal isn’t a rate limit. When the upstream refuses because of its own budget or quota, Inferrail returns INFERRAIL_E014 (HTTP 402). That covers LiteLLM budget_exceeded (429), Vercel quota_for_entity_exceeded (402), OpenAI insufficient_quota (429), a 402, and a 403/429 whose detail/message mentions a budget (otari: “API key has exceeded budget limit”). The request is not retried, and the details carry the upstream status and type. We return 402 rather than passing through a 429 because SDKs automatically retry 429s. The upstream’s free text is used only to classify, and is never stored. The reservation is released (an HTTP refusal isn’t billed).

Consequences