Accepted (pending review)
ADR-0021 made block-mode admission atomic. But protecting a single agent
run still means creating a budget object for its work_id first
(inferrail budget set --scope work_id --scope-value <id> …), and doing
that again for every new run. Most runs get their id when they start, so
this setup step is the main friction in “give this run a dollar ceiling”.
The desired flow:
mark the work (work_id)
→ declare a budget, or inherit a default
→ no separate setup
→ enforced at admission
→ a final per-run record
A request that carries X-Inferrail-Attribute-Work-Id: <id> can be
covered by a per_work, block-mode budget for that id in three ways:
work_id:<id>:per_work
(created with the CLI or local API, or created earlier by 2 or 3
below).X-Inferrail-Budget-Usd: <amount>,
and no budget exists yet for that id.budgets.per_work_default_usd, no
budget exists yet for that id, and the request doesn’t declare one.With 2 or 3, the budget row is created inside the same admission
transaction that reserves the request (ADR-0021’s BEGIN IMMEDIATE).
So when concurrent first requests for a new id all declare the same
budget, exactly one creates it, and every one of them is admitted
against it. No request for an unseen id can slip past before the budget
exists.
A request with no work_id isn’t affected. Neither is a work_id
request when no budget is stored, declared, or defaulted. Global and
project budgets still apply exactly as before.
INFERRAIL_E013, HTTP 400) before any provider call. Silently
applying either value would hide a bug. An identical declaration
(retries, or every call of the run sending the same header) is
accepted.budgets.per_work_max_usd, if set, is the most a
header may declare. A larger declaration is refused (E013), never
clamped.budgets.allow_declared_budgets: false. The header is then refused
(E013), never silently ignored.INFERRAIL_GATEWAY_TOKEN is set, only authenticated callers get this
far. A declaration only ever adds a limit for an id that has no
budget yet, so a caller can’t use it to raise anyone’s limit. The worst
a hostile caller can do is declare a small budget for a work_id
before its real owner does, which refuses the owner’s later calls with
a clear conflict error. That is a denial of service by someone who is
already trusted to call the gateway, not overspend. Operators who need
more should keep declarations off and store budgets themselves.gateway/attribution.py).work_id. A sub-agent with its own work_id gets its own budget,
and its spend does not also count against the parent’s. There is
no budget hierarchy in this version; a project budget can bound a
group of runs.work_id, a value above the operator ceiling, a conflict, or
declarations turned off.(scope, scope_value) index. Measured: median admission stayed at
2–3 ms with 0 to 50,000 accumulated per-run rows (it was 914 ms at
50,000 with a full scan). inferrail budget rm removes one row;
lifecycle cleanup (e.g. expiring finished runs) is a follow-up.A team can keep its current OpenAI-compatible gateway and add Inferrail in front of it just for per-run control. This was tested locally in front of LiteLLM and otari; other gateways haven’t been tested:
framework → Inferrail → existing gateway → provider
Two opt-in settings on an openai_compatible provider remove the
friction found in testing:
price_as: openai (or anthropic) is the operator asserting that this
upstream bills at that vendor’s list prices. The built-in catalog then
applies, also for vendor-prefixed ids (openai/gpt-4o-mini,
openai:gpt-4o-mini). The price’s recorded source says it was
applied through price_as, so it’s never presented as verified.
Explicit pricing: overrides still win. Without either, a block budget
refuses the model (E012), as before.request_stream_usage: true asks the upstream for the final usage
chunk on streams, as Inferrail already does for verified OpenAI.
Without it, a stream whose client didn’t request usage stays unpriced
and its reservation is held.Each gateway enforces only its own budgets, so nothing is counted twice.
A downstream budget refusal isn’t a rate limit. When the upstream
refuses because of its own budget or quota, Inferrail returns
INFERRAIL_E014 (HTTP 402). That covers LiteLLM budget_exceeded (429),
Vercel quota_for_entity_exceeded (402), OpenAI insufficient_quota
(429), a 402, and a 403/429 whose detail/message mentions a budget
(otari: “API key has exceeded budget limit”). The request is not
retried, and the details carry the upstream status and type. We return
402 rather than passing through a 429 because SDKs automatically retry
429s. The upstream’s free text is used only to classify, and is never
stored. The reservation is released (an HTTP refusal isn’t billed).
inferrail budget list, same reservation semantics).budgets.enabled: true (a sqlite receipts store),
as all budgets do.