An agent run rarely makes one model call. One user request becomes a plan, then several tool calls, often in parallel, then a summary. This recipe puts a dollar ceiling on that whole run, enforced outside the agent, and gives you the run’s final cost afterwards.
What you get:
pip install inferrail
export OPENAI_API_KEY=sk-...
In your app:
import inferrail
base_url = inferrail.start() # e.g. http://127.0.0.1:53721/v1
inferrail.start() runs the Inferrail gateway on a background thread in
your process, on a free local port, with budgets on and no config file.
It reads the provider key from your environment. Receipts and budgets
go to SQLite files in your user data directory, so inferrail work
finds them later. Calling it again returns the same URL;
inferrail.stop() shuts it down (it also stops when your process
exits).
Or run it as its own process, for example one gateway shared by several services:
inferrail.yaml:
providers:
openai: {type: openai, api_key_env: OPENAI_API_KEY}
default_provider: openai # each request's model id is passed through
receipts: {sink: sqlite, path: ./receipts.db}
budgets: {enabled: true, path: ./budgets.db}
inferrail serve --config inferrail.yaml # http://127.0.0.1:8000/v1
Budgets and admission work the same either way. The snippets below use
base_url; with inferrail serve it’s http://127.0.0.1:8000/v1.
Point your OpenAI client at Inferrail and send two headers on every call of the run:
from openai import AsyncOpenAI
run_id = "run-7f3a" # your run / job / request id
client = AsyncOpenAI(
base_url=base_url, # from inferrail.start(), or your gateway's URL
api_key="unused", # the provider key lives in the gateway
default_headers={
"X-Inferrail-Attribute-Work-Id": run_id, # which run this call belongs to
"X-Inferrail-Budget-Usd": "0.05", # the run's dollar budget
},
)
The first request of a run creates its budget. Every later request of the same run must declare the same amount (retries do this naturally).
import asyncio
from openai import APIStatusError
async def one_call(i):
try:
await client.chat.completions.create(
model="gpt-4o-mini", max_tokens=200,
messages=[{"role": "user", "content": f"Summarize item {i}."}],
)
return "answered"
except APIStatusError as e:
if e.status_code == 402: # refused before the provider
return "refused"
raise
outcomes = await asyncio.gather(*(one_call(i) for i in range(8)))
When the run’s remaining budget can’t cover the next call’s reservation, that call gets a 402 like:
{"error": {"code": "INFERRAIL_E010",
"message": "budget 'work_id:run-7f3a:per_work' ... would be exceeded ...",
"details": {"budget_id": "work_id:run-7f3a:per_work", "limit_usd": "0.05",
"spent_so_far_usd": "...", "reserved_usd": "...", "...": "..."}}}
A complete, runnable version is in
examples/agent_run_budget.py. It
is exercised in CI against the gateway: 8 parallel calls on a budget with
room for 4 → 4 answered, 4 refused, 4 provider calls.
inferrail work run-7f3a # with inferrail serve: add --config inferrail.yaml
Work: run-7f3a
Inference receipts: 8
Known attributed inference cost: $0.000...
Unknown-cost inference receipts: 0
Refused calls appear as receipts with no cost. Nothing was billed for them.
Every framework below talks to Inferrail through its normal OpenAI-compatible
client. Short standalone pages: LangGraph,
OpenAI Agents SDK,
CrewAI, Strands Agents,
Vercel AI SDK. What changes is only where the run id and budget go. Each
snippet uses the base_url from step 1. Each was run with two concurrent
runs, both through inferrail serve (0.4.8) and through
inferrail.start() (the TypeScript ones only through inferrail serve):
the runs stayed separate, and the run with the small budget got a 402.
gpt-4o-mini in the snippets is only an example. Inferrail doesn’t choose
a model: send any model id your account can use, and it’s passed to the
provider as-is. A dollar budget needs a price for that model.
inferrail models lists your provider’s models and which ones have a
price. For a model Inferrail has no price for yet, pass one yourself:
base_url = inferrail.start(pricing={"openai": {"<model id>": {
"input_usd_per_million": "...", "output_usd_per_million": "...",
"source": "OpenAI pricing page", "verified_date": "2026-10-01",
}}})
Without a price, that model still works, but a run with a budget refuses
it (INFERRAIL_E012) rather than guess what it costs.
One model object; pass the run’s headers on each call:
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(model="gpt-4o-mini", base_url=base_url,
api_key="unused", max_tokens=200)
llm.invoke("Summarize the ticket.", extra_headers={
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
})
Tested with langchain-openai 1.6.7.
An agent makes several model calls per run, so give each run its own model with the headers set once:
from langchain.agents import create_agent
from langchain_openai import ChatOpenAI
def agent_for_run(run_id: str, budget_usd: str):
model = ChatOpenAI(
model="gpt-4o-mini", base_url=base_url,
api_key="unused", max_tokens=200,
default_headers={
"X-Inferrail-Attribute-Work-Id": run_id,
"X-Inferrail-Budget-Usd": budget_usd,
},
)
return create_agent(model, tools=[...])
agent_for_run("run-7f3a", "0.50").invoke(
{"messages": [{"role": "user", "content": "..."}]})
A refused call raises openai.APIStatusError with status_code == 402.
Tested with langchain 1.4.3 and langgraph 1.2.12
(create_react_agent behaves the same).
One client for the whole process; the run’s headers go in its
RunConfig, so concurrent runs don’t share a budget by accident:
from agents import Agent, ModelSettings, RunConfig, Runner, set_default_openai_api, set_default_openai_client
from openai import AsyncOpenAI
set_default_openai_api("chat_completions")
set_default_openai_client(AsyncOpenAI(base_url=base_url, api_key="unused"))
agent = Agent(name="support", instructions="...", model="gpt-4o-mini",
model_settings=ModelSettings(max_tokens=200))
await Runner.run(agent, "...", run_config=RunConfig(model_settings=ModelSettings(extra_headers={
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
})))
Tested with openai-agents 0.22.3 (Chat Completions API; the Responses
API isn’t supported by Inferrail yet).
OpenAIChatGenerator with api_base_url pointing at Inferrail, and the
run’s headers in generation_kwargs:
pipeline.run({
"llm": {"generation_kwargs": {"max_tokens": 300, "extra_headers": {
"X-Inferrail-Attribute-Work-Id": "support-ticket-4812",
"X-Inferrail-Budget-Usd": "0.05",
}}},
# ... your other components' inputs
})
A refused call raises openai.APIStatusError (status_code == 402) from
the generator; inside pipeline.run(...) Haystack wraps it in
PipelineRuntimeError, with the 402 as its __cause__. Tested with
haystack-ai 3.2.0.
Give each crew run its own LLM with the run’s headers. Every agent in
the crew that uses it shares the run’s budget, including delegated work:
from crewai import LLM
llm = LLM(model="gpt-4o-mini", base_url=base_url,
api_key="unused", max_tokens=200,
extra_headers={
"X-Inferrail-Attribute-Work-Id": "crew-run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
})
# pass llm=llm to each Agent in this crew run
CrewAI retries a failed call up to its retry limit; each retry of a
refused call is also refused before the provider, so it costs nothing.
Use inferrail 0.4.10 or later in the same environment as CrewAI
(earlier versions required an mcp version CrewAI can’t install).
Tested with crewai 1.15.23 and inferrail 0.4.10.
from llama_index.llms.openai_like import OpenAILike
llm = OpenAILike(model="gpt-4o-mini", api_base=base_url,
api_key="unused", is_chat_model=True, max_tokens=200,
default_headers={
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
})
Create one per run. Tested with llama-index-llms-openai-like 0.8.0.
Use the Chat Completions client (Inferrail doesn’t serve the Responses API yet), one per run:
from agent_framework import Agent
from agent_framework.openai import OpenAIChatCompletionClient
client = OpenAIChatCompletionClient(
model="gpt-4o-mini", base_url=base_url, api_key="unused",
default_headers={
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
})
agent = Agent(client=client, instructions="...", tools=[...],
default_options={"max_tokens": 200})
A refused call surfaces as ChatClientException wrapping the 402.
Tested with agent-framework-core 1.19.0 and agent-framework-openai
1.14.4.
OpenAIModel passes client_args to the OpenAI client, so the run’s
headers go in default_headers. Build one model per run:
from strands import Agent
from strands.models.openai import OpenAIModel
def agent_for_run(run_id: str, budget_usd: str):
model = OpenAIModel(
client_args={
"base_url": base_url,
"api_key": "unused",
"default_headers": {
"X-Inferrail-Attribute-Work-Id": run_id,
"X-Inferrail-Budget-Usd": budget_usd,
},
},
model_id="gpt-4o-mini",
params={"max_tokens": 200},
)
return Agent(model=model, tools=[...])
agent_for_run("run-7f3a", "0.50")("...")
A refused call ends the agent loop with an EventLoopException wrapping
the 402. Tested with strands-agents 1.57.1.
Use ADK’s LiteLlm model with api_base pointing at Inferrail and the
run’s headers in extra_headers, one per run:
from google.adk.agents import LlmAgent
from google.adk.models.lite_llm import LiteLlm
def agent_for_run(run_id: str, budget_usd: str):
model = LiteLlm(
model="openai/gpt-4o-mini", api_base=base_url,
api_key="unused", max_tokens=200,
extra_headers={
"X-Inferrail-Attribute-Work-Id": run_id,
"X-Inferrail-Budget-Usd": budget_usd,
},
)
return LlmAgent(name="support", model=model, instruction="...", tools=[...])
In ADK 2.x a refused call arrives as an event with error_code set
rather than an exception. Stop reading events at that one: the runner
raises if you keep iterating. Tested with
google-adk 2.10.0 (and 1.10.0) and litellm 1.103.2. Only Chat
Completions models routed through LiteLlm are covered, not native
Gemini.
OpenAIServerModel forwards client_kwargs to the OpenAI client:
from smolagents import OpenAIServerModel, ToolCallingAgent
model = OpenAIServerModel(
model_id="gpt-4o-mini", api_base=base_url,
api_key="unused", max_tokens=200,
client_kwargs={"default_headers": {
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
}},
)
agent = ToolCallingAgent(tools=[...], model=model)
Create one model per run. A refused call raises AgentGenerationError
wrapping the 402. Tested with smolagents 1.26.0.
dspy.LM passes extra_headers through to the request. Scope the LM to
one program run with dspy.context:
import dspy
lm = dspy.LM("openai/gpt-4o-mini", api_base=base_url,
api_key="unused", max_tokens=200, cache=False,
extra_headers={
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
})
with dspy.context(lm=lm):
program(question="...")
A refused call raises LMBillingError. Turn off DSPy’s cache
(cache=False) if you want every call to reach the gateway and count.
Tested with dspy 3.4.0.
from agno.agent import Agent
from agno.models.openai.like import OpenAILike
model = OpenAILike(id="gpt-4o-mini", base_url=base_url,
api_key="unused", max_tokens=200,
default_headers={
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
})
agent = Agent(model=model, tools=[...])
run = agent.run("...")
Create one model per run. Agno doesn’t raise on a refused call: the run
comes back with run.status == RunStatus.error and the budget message as
its content, so check the status. Tested with agno 3.1.0.
AG2 1.x takes an OpenAIConfig with default_headers, one per run:
from ag2 import Agent
from ag2.config.openai.config import OpenAIConfig
def agent_for_run(run_id: str, budget_usd: str) -> Agent:
config = OpenAIConfig(
model="gpt-4o-mini", base_url=base_url, api_key="unused", max_tokens=200,
default_headers={
"X-Inferrail-Attribute-Work-Id": run_id,
"X-Inferrail-Budget-Usd": budget_usd,
},
)
return Agent("support", "...", config=config)
reply = await agent_for_run("run-7f3a", "0.50").ask("...")
A refused call raises openai.APIStatusError (status_code == 402).
Tested with ag2 1.1.1.
from agentscope.credential import OpenAICredential
from agentscope.model import OpenAIChatModel
model = OpenAIChatModel(
credential=OpenAICredential(api_key="unused", base_url=base_url),
model="gpt-4o-mini",
parameters=OpenAIChatModel.Parameters(max_tokens=200),
client_kwargs={"default_headers": {
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
}},
)
Create one model per run. A refused call raises openai.APIStatusError
(status_code == 402). Needs inferrail 0.4.10 or later in the same
environment. Tested with agentscope 2.0.9.
from camel.agents import ChatAgent
from camel.models import ModelFactory
from camel.types import ModelPlatformType
model = ModelFactory.create(
model_platform=ModelPlatformType.OPENAI_COMPATIBLE_MODEL,
model_type="gpt-4o-mini", url=base_url, api_key="unused",
model_config_dict={"max_tokens": 200, "extra_headers": {
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
}},
)
agent = ChatAgent(model=model)
Create one model per run (or per Workforce run). A refused call raises
openai.APIStatusError (status_code == 402). camel-ai 0.2.90 fails to import with mcp 2.x, so keep mcp<2 in
that environment (inferrail 0.4.10 or later doesn’t require mcp).
Tested with camel-ai 0.2.90, mcp 1.29 and inferrail 0.4.10.
Give the connector an AsyncOpenAI client with the run’s headers:
from openai import AsyncOpenAI
from semantic_kernel.connectors.ai.open_ai import OpenAIChatCompletion
service = OpenAIChatCompletion(
ai_model_id="gpt-4o-mini",
async_client=AsyncOpenAI(base_url=base_url, api_key="unused", default_headers={
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
}),
)
Create one service per run. A refused call raises
ServiceResponseException wrapping the 402. Needs inferrail 0.4.10 or later in
the same environment. Tested with semantic-kernel 1.44.1.
import langroid as lr
import langroid.language_models as lm
llm_config = lm.OpenAIGPTConfig(
chat_model="gpt-4o-mini", api_base=base_url, api_key="unused",
max_output_tokens=200,
headers={
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
},
)
agent = lr.ChatAgent(lr.ChatAgentConfig(llm=llm_config))
Create one config per run. Langroid retries a refused call several times
before raising; each retry is refused before the provider, so retries
cost nothing. Tested with langroid 0.68.2.
from beeai_framework.adapters.openai import OpenAIChatModel
model = OpenAIChatModel(
"gpt-4o-mini", base_url=base_url, api_key="unused",
settings={"extra_headers": {
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
}},
)
Create one model per run. A refused call raises ChatModelError. Tested
with beeai-framework 0.1.85.
from griptape.artifacts import ErrorArtifact
from griptape.drivers.prompt.openai import OpenAiChatPromptDriver
from griptape.structures import Agent
from openai import OpenAI
driver = OpenAiChatPromptDriver(
model="gpt-4o-mini", base_url=base_url, api_key="unused", max_tokens=200,
client=OpenAI(base_url=base_url, api_key="unused", default_headers={
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
}),
)
result = Agent(prompt_driver=driver).run("...")
if isinstance(result.output, ErrorArtifact):
... # the run reached its budget
Create one driver per run. Griptape doesn’t raise on a refused call: the
task finishes with an ErrorArtifact carrying the 402, so check the
output. Tested with griptape 1.13.0.
Atomic Agents uses an instructor client, so the headers go on the
OpenAI client underneath:
import instructor
from openai import OpenAI
from atomic_agents import AgentConfig, AtomicAgent, BasicChatInputSchema, BasicChatOutputSchema
client = instructor.from_openai(OpenAI(base_url=base_url, api_key="unused", default_headers={
"X-Inferrail-Attribute-Work-Id": "run-7f3a",
"X-Inferrail-Budget-Usd": "0.50",
}))
agent = AtomicAgent[BasicChatInputSchema, BasicChatOutputSchema](
config=AgentConfig(client=client, model="gpt-4o-mini", model_api_parameters={"max_tokens": 200}))
Create one client per run. instructor retries a failed call; each retry
of a refused call is refused before the provider. A refused run raises
InstructorRetryException with the 402 in it. Tested with
atomic-agents 2.10.3.
Inferrail is a separate process, so a TypeScript app can use it over HTTP. Use the OpenAI-compatible provider and pass the run’s headers per call:
import { createOpenAICompatible } from '@ai-sdk/openai-compatible';
import { generateText } from 'ai';
const inferrail = createOpenAICompatible({
name: 'inferrail',
baseURL: 'http://127.0.0.1:8000/v1',
apiKey: 'unused',
includeUsage: true,
});
await generateText({
model: inferrail('gpt-4o-mini'),
maxOutputTokens: 200,
headers: {
'X-Inferrail-Attribute-Work-Id': 'run-7f3a',
'X-Inferrail-Budget-Usd': '0.50',
},
tools: { /* ... */ },
prompt: '...',
});
Every step of a multi-step generateText call carries the same headers.
A refused call throws an error with statusCode === 402. Tested with
ai 7.0.126 and @ai-sdk/openai-compatible 3.0.62. A TypeScript app
can’t call inferrail.start(), so run the gateway next to it with
pip install inferrail and inferrail serve --quickstart --app-mode
(no config file; inferrail work <run-id> reads its receipts).
Mastra agents take an AI SDK model, so the OpenAI-compatible provider works as-is. Keep one agent and pass the run’s headers on each call:
import { Agent } from '@mastra/core/agent';
import { createOpenAICompatible } from '@ai-sdk/openai-compatible';
const inferrail = createOpenAICompatible({
name: 'inferrail', baseURL: 'http://127.0.0.1:8000/v1', apiKey: 'unused', includeUsage: true,
});
const agent = new Agent({
id: 'support', name: 'support', instructions: '...',
model: inferrail.chatModel('gpt-4o-mini'), tools: { /* ... */ },
});
await agent.generate('...', { modelSettings: {
maxOutputTokens: 200,
headers: { 'X-Inferrail-Attribute-Work-Id': 'run-7f3a', 'X-Inferrail-Budget-Usd': '0.50' },
} });
A refused call throws an error with statusCode === 402. Run the gateway
next to the app (inferrail serve --quickstart --app-mode). Tested with
@mastra/core 1.74.0 through inferrail serve.
VoltAgent agents also take an AI SDK model. Keep one agent and pass the run’s headers on each call:
import { Agent } from '@voltagent/core';
import { createOpenAICompatible } from '@ai-sdk/openai-compatible';
const inferrail = createOpenAICompatible({
name: 'inferrail', baseURL: 'http://127.0.0.1:8000/v1', apiKey: 'unused', includeUsage: true,
});
const agent = new Agent({
name: 'support', instructions: '...',
model: inferrail.chatModel('gpt-4o-mini'), tools: [ /* ... */ ],
});
await agent.generateText('...', {
maxOutputTokens: 200,
headers: { 'X-Inferrail-Attribute-Work-Id': 'run-7f3a', 'X-Inferrail-Budget-Usd': '0.50' },
});
A refused call throws an error with statusCode === 402, and VoltAgent
doesn’t retry it. Run the gateway next to the app
(inferrail serve --quickstart --app-mode). Tested with
@voltagent/core 2.11.0, ai 6.0.300 and @ai-sdk/openai-compatible
2.0.81 through inferrail serve.
@openai/agents has no per-call header setting, so give each run its
own Runner whose OpenAI client carries the run’s headers. The agent
itself is shared:
import { Agent, Runner, OpenAIProvider } from '@openai/agents';
import OpenAI from 'openai';
const agent = new Agent({
name: 'support', instructions: '...', model: 'gpt-4o-mini',
modelSettings: { maxTokens: 200 },
});
function runnerFor(runId: string, budgetUsd: string) {
const client = new OpenAI({
baseURL: 'http://127.0.0.1:8000/v1', apiKey: 'unused',
defaultHeaders: { 'X-Inferrail-Attribute-Work-Id': runId, 'X-Inferrail-Budget-Usd': budgetUsd },
});
return new Runner({ modelProvider: new OpenAIProvider({ openAIClient: client, useResponses: false }) });
}
await runnerFor('run-7f3a', '0.50').run(agent, '...');
useResponses: false matters: Inferrail serves Chat Completions, not the
Responses API. A refused call throws an error with status === 402; the
OpenAI client doesn’t retry it. Run the gateway next to the app
(inferrail serve --quickstart --app-mode). Tested with @openai/agents
0.18.0 through inferrail serve.
Inferrail can sit in front of an OpenAI-compatible gateway you already run, just for per-run budgets:
providers:
gw:
type: openai_compatible
api_key_env: GATEWAY_KEY # e.g. your LiteLLM virtual key
base_url: http://127.0.0.1:4000/v1
price_as: openai # you assert the gateway bills at OpenAI list prices
request_stream_usage: true # ask the gateway for stream usage
This was tested locally in front of LiteLLM and otari. If that gateway
refuses a call because of its own budget, you get INFERRAIL_E014 (HTTP
402, not retried). See ADR 0022.
Use the config with inferrail serve --config inferrail.yaml, or in
process with inferrail.start("inferrail.yaml").
Keep your OpenRouter key and models, and add a per-run ceiling:
providers:
openrouter:
type: openai_compatible
api_key_env: OPENROUTER_API_KEY
base_url: https://openrouter.ai/api/v1
price_as: openai # prices openai/* models at OpenAI list prices (your assertion)
request_stream_usage: true
default_provider: openrouter # send OpenRouter model ids as-is, e.g. "openai/<model>"
receipts: {sink: sqlite, path: ./receipts.db}
budgets: {enabled: true, path: ./budgets.db}
price_as: openai only prices openai/... model ids, and only those
whose name is in Inferrail’s OpenAI price list. For any other OpenRouter
model, add a pricing: entry with its price, or a budgeted run refuses it
(INFERRAIL_E012) instead of guessing. inferrail models --config
inferrail.yaml lists OpenRouter’s models and which ones have a price.
Checked against the live OpenRouter API: a run’s over-budget call is
refused before any request reaches OpenRouter, OpenRouter’s own “out of
credits” 402 comes back as INFERRAIL_E014 (not retried), no
X-Inferrail-* header is sent to OpenRouter, and receipts hold no
prompts or responses. A priced, answered call was checked against a local
OpenRouter-shaped stand-in, not yet against the live API.
max_tokens (or
max_completion_tokens) at list price. Set max_tokens: without
it, Inferrail assumes 4,096 output tokens, which can refuse calls far
earlier than their real cost would. If a call’s actual cost exceeds its
reservation, it still completes, the excess is recorded as
budget_overrun_usd, and later calls are refused. This is a ceiling on
admission, not a guarantee that spend can never pass the limit.budget_held_usd, never as cost.INFERRAIL_E012). Add
a pricing: override to use one.INFERRAIL_GATEWAY_TOKEN,
budgets.per_work_max_usd, or budgets.allow_declared_budgets: false
to control this.Did this work for you, or not? Open an issue with what happened. It directly shapes what we fix next.