A per-run budget usually works like this: before each model call, read how much the run has spent, compare it with the limit, then make the call and record its cost when the response comes back. That is exact when calls run one at a time. It isn’t when they run at once.
Every call that starts before the first one finishes reads the same “spent so far”. Parallel tool calls, sub-agents working at the same time, a retry racing the original, or a batch job with a thread pool can all pass the check together. The cost of a call isn’t known until its response ends, so the window isn’t a few milliseconds: it’s the whole duration of the call.
With a budget worth 10 calls and 30 calls started together, a check-then-record budget lets all 30 through.
limit - committed - outstanding reservations >= estimate,
and record the reservation in that same transaction. The estimate is the
input tokens plus max_tokens at the model’s price, which is why
budgets need max_tokens set.The per-run budget benchmark
is a small reproducible test: a fake OpenAI-compatible upstream, a budget
worth exactly 10 calls at true cost, then 30 calls at once (plus 5 more), 15
calls one at a time, and the same with streams. It counts how many calls
reached the upstream. Point burst.py at any OpenAI-compatible endpoint
that enforces a budget:
uvicorn fake_upstream:app --port 9400
python burst.py --url http://127.0.0.1:<port>/v1/chat/completions --nonce --n 30
It has already been useful outside this project: it showed that one proxy’s per-run cap let a 30-call burst through, the proxy’s maintainer traced it to check-then-record admission and fixed it by reserving at admission, and the rerun now holds at 10. The results table lists the exact versions and configs tested, including an earlier Inferrail release that had the same problem.
Inferrail implements the steps above as an OpenAI- and Anthropic-compatible gateway that runs in your process. The run’s id and budget are two request headers, the reservation happens in one SQLite transaction, and a refused call gets HTTP 402 before it reaches the provider:
import inferrail
from openai import OpenAI
base_url = inferrail.start()
client = OpenAI(base_url=base_url, api_key="unused", default_headers={
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
})
Framework-specific setup (LangGraph, CrewAI, the OpenAI Agents SDK, Strands, ADK, DSPy, the Vercel AI SDK and more) is in Give one AI agent run a dollar budget.