How to send traffic through a self-hosted Inferrail gateway, attach attribution, and group requests into units of work. For the exact supported request surface, see PRODUCT.md.
The gateway process calls the provider, so the gateway process needs
the provider key. A key that exists only in your application’s process
is not passed along. Set the key in the terminal where you run
inferrail serve:
export OPENAI_API_KEY=... # for /v1/chat/completions
export ANTHROPIC_API_KEY=... # for /v1/messages
inferrail serve --quickstart
--quickstart registers both providers and passes any model id through
to the matching one. Only the provider whose key is set will succeed.
Real requests are billed by your provider as usual.
The model ids in the examples below are only examples. Inferrail doesn’t
choose a model: send any model your account can use. inferrail models
lists them and shows which ones have a price (needed under a dollar
budget).
Clients then point at the gateway. Unless you set
INFERRAIL_GATEWAY_TOKEN, the gateway ignores the client’s API key, so
any placeholder works.
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="not-needed")
client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Say hello in five words."}],
extra_headers={"X-Inferrail-Attribute-Customer": "acme"},
)
Or, without code changes: export OPENAI_BASE_URL=http://127.0.0.1:8000/v1.
With curl:
curl http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "X-Inferrail-Attribute-Customer: acme" \
-d '{"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Say hello in five words."}]}'
The response is the standard choices/usage shape plus a non-standard
inferrail block (route, provider, latency, retries) that OpenAI clients
ignore. Supported: text messages, streaming (stream: true), and
tool/function calling, structured outputs (response_format), and the
other provider-valid fields listed in PRODUCT.md. Rejected
with an error rather than silently dropped: n != 1, non-text content
parts, fields whose cost the gateway can’t account for, and any unknown
field. See examples/basic_chat_request.py.
POST /v1/messages is a separate Anthropic-compatible passthrough with
streaming and tool use (ADR 0014).
The Anthropic SDKs append /v1/messages themselves, so their base URL
stops at the origin, without /v1:
import anthropic
client = anthropic.Anthropic(base_url="http://127.0.0.1:8000", api_key="not-needed")
client.messages.create(
model="claude-haiku-4-5-20251001",
max_tokens=256,
messages=[{"role": "user", "content": "Say hello in five words."}],
)
Or export ANTHROPIC_BASE_URL=http://127.0.0.1:8000. See
examples/anthropic_messages_request.py.
Claude Code is not supported yet. Current Claude Code versions send
request fields the gateway does not forward (thinking,
context_management, output_config), so the gateway rejects the first
request with HTTP 400.
The
inferrail serve --quickstartbanner in release 0.4.3 prints the Anthropic base URL with a trailing/v1, which makes the SDK request/v1/v1/messages. Use the origin shown above. The banner is fixed onmain.
"model" first selects a named route from inferrail.yaml (for example
default), which maps to a provider and model. If default_provider (or
default_anthropic_provider) is set, a model that matches no route is
forwarded unchanged to that provider. Quickstart turns this passthrough
on; explicit configs leave it off unless you set it. Named routes always
win. Design: ADR 0007.
Cost is computed only when the provider reports usage and a verified
price is on file for that provider and model. Otherwise pricing and
estimated_cost_usd are null, never a guessed 0. The built-in price
catalog applies only to the providers’ own default endpoints; for a
custom base_url, add prices under pricing: in inferrail.yaml.
These are configuration examples. CI tests the OpenAI wire protocol these frameworks use (test_agent_e2e.py), not the frameworks themselves.
# LangChain
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
base_url="http://127.0.0.1:8000/v1",
api_key="not-needed", # or your INFERRAIL_GATEWAY_TOKEN if auth is enabled
model="default",
)
# LlamaIndex
from llama_index.llms.openai_like import OpenAILike
llm = OpenAILike(
model="default",
api_base="http://127.0.0.1:8000/v1",
api_key="not-needed",
is_chat_model=True,
context_window=8192,
)
# CrewAI
from crewai import LLM
llm = LLM(
model="openai/default", # "openai/" prefix required by CrewAI
base_url="http://127.0.0.1:8000/v1",
api_key="not-needed",
)
Inferrail has no native voice support. It does not handle audio, speech-to-text, text-to-speech, the OpenAI Realtime API, or WebSocket sessions, and it does not account for full call cost.
A voice stack can still route its text LLM stage through Inferrail if that stage lets you set an OpenAI- or Anthropic-compatible base URL and sends a supported request shape (text messages, optionally streaming or tool calls). Only that stage’s tokens and cost appear in receipts.
No voice framework integration (LiveKit, Pipecat, Vapi, Retell, or
others) has been tested by this project. Treat compatibility as
something to verify in your own stack: send one request, then check
that a receipt appears in inferrail report.
Three ways to attach business context, all landing in the same
attributes: dict[str, str] on the receipt:
X-Inferrail-Attribute-<Name>: <value>, for
example X-Inferrail-Attribute-Task-Id: bug_9281. These headers are
never forwarded to the provider (attribution.py).inferrail try): --customer, --workflow, or generic
-a <name>=<value>.inferrail.track_task attaches
X-Inferrail-Attribute-Task-Id to every request inside a with block
or decorated function.import inferrail
from openai import OpenAI
client = OpenAI(
base_url="http://127.0.0.1:8000/v1",
api_key="not-needed",
# base_url must match the client's own base_url: the header is only
# attached to requests going to that destination.
http_client=inferrail.attributed_http_client(base_url="http://127.0.0.1:8000/v1"),
)
@inferrail.track_task(task_id="bug_9281")
def fix_bug():
client.chat.completions.create(...) # tagged automatically
run_subagent() # nested calls too
Sync and async are both supported (attributed_async_http_client), and
concurrent tasks never cross-contaminate. See
ADR 0009.
Attribute values are stored exactly as sent. Use identifiers, not names, emails, secrets, or message text.
Then slice spend by any dimension:
inferrail report --by customer
inferrail report --by workflow
inferrail report --by provider # also: model, route, or any attribute name
Give related requests the same work_id, then declare an outcome when
your application knows one:
inferrail try "Review this contract clause" --model <model id> -a work_id=contract_review_42
inferrail try "Identify remaining risks" --model <model id> -a work_id=contract_review_42
inferrail work outcome contract_review_42 --status completed
inferrail work contract_review_42
inferrail work --all
inferrail try sends one real, billed request with OPENAI_API_KEY to
the model you name (inferrail models lists yours). Over
HTTP, use X-Inferrail-Attribute-Work-Id: contract_review_42.
Work Economics reports the known attributed inference cost for that work. If any receipt in the work has unknown cost, the report counts it separately rather than folding it into the total, so a known subtotal is not a complete bill. It is not COGS, margin, or business value, and Inferrail does not interpret your outcome labels.
inferrail transaction <task-id> gives the older receipt-only grouping
by task_id (ADR 0008).
inferrail mcp is a stdio MCP server with two read-only tools over your
local receipts file:
| Tool | What it answers |
|---|---|
get_spend |
Known cost, tokens, and request counts grouped by provider, model, route, or any attribute you tag requests with (customer, workflow, work_id), optionally within a time window. Requests with unknown pricing are counted separately, not as $0. |
get_health |
Whether the gateway answers GET /health, plus the most recent receipt. |
Neither runs inference, spends provider budget, changes configuration, or
writes files. Receipts store usage and cost metadata without persisting
prompt or response bodies, so the tools have none to return. Grouping by
customer, workflow, or work_id only covers requests that were sent
with that tag (Attribution).
Claude Code:
claude mcp add inferrail -e INFERRAIL_RECEIPTS_PATH=/absolute/path/to/inferrail-receipts.jsonl -- uvx --with "mcp>=2.0" inferrail mcp
Claude Desktop, Cursor, and other clients that use mcpServers (VS Code
uses the same entry under servers):
{
"mcpServers": {
"inferrail": {
"command": "uvx",
"args": ["inferrail", "mcp"],
"env": {
"INFERRAIL_RECEIPTS_PATH": "/absolute/path/to/inferrail-receipts.jsonl"
}
}
}
}
Without uvx: pip install inferrail, then use inferrail mcp as the
command.
Set INFERRAIL_RECEIPTS_PATH to your receipts file. Clients start the
server from their own working directory, so the default
./inferrail-receipts.jsonl is rarely the right place. For
serve --app-mode, point it at receipts.db in Inferrail’s data
directory (~/.local/share/inferrail on Linux,
~/Library/Application Support/inferrail on macOS, %APPDATA%\inferrail
on Windows).
Then ask, for example: “How much did work contract-review-42 cost?” If
your requests carried work_id=contract-review-42, the agent calls
get_spend with by: "work_id" and reads that group. Full tool
contract: inferrail-mcp/README.md.