For agents and the people who run them
Inferrail HostedOne budget per run, across your whole fleet.
Give each AI run a dollar ceiling in a request header. Every call in that run shares it, including parallel calls from different agents and machines, and a call that would push the run past it is refused before it reaches OpenAI or Anthropic. You bring your own provider key with each request; we never store it.
Experimental · 2,000 governed runs a month free · then $0.001 per run · no subscription
Not open yet. This page describes the terms Inferrail Hosted will launch with. The hosted service isn't accepting workspaces yet; until it is, the open-source gateway does the same per-run enforcement on your own machine.
Price
A governed run: a distinct work id (X-Inferrail-Attribute-Work-Id) in a calendar month (UTC). Calls inside a run are not charged separately.
At least one call in the run is answered by your provider. A run whose calls all fail, or are all refused by its own budget, costs nothing.
2,000 governed runs per workspace per month.
Prepaid credits at $0.001 per run. People pay by card: $10 for 10,000 runs. Agents pay in USDC over x402: $1 for 1,000 runs, with no account.
Credits are prepaid. No subscription, no auto top-up, no overage bill. When the allowance and credits run out, calls are refused before reaching your provider, with a link to buy more.
Billed by your provider on your own key. Inferrail doesn't resell model access or mark it up.
The open-source gateway is free and does everything on one machine: pip install inferrail. Hosted is for when many agents and machines must share one budget ledger and you'd rather not run it.
Start
POST /v1/workspacesreturns a workspace key, once. No sign-up form.- Point an OpenAI or Anthropic SDK at
/v1with the workspace key as the API key, and send your provider key inX-Provider-Api-Key. - Name the run and set its ceiling:
X-Inferrail-Attribute-Work-Id,X-Inferrail-Budget-Usd. GET /v1/work/<id>shows what the run cost.GET /pricingand/llms.txtpublish these terms for agents.
Past your allowance with no credits left, a call gets HTTP 402 allowance_exhausted, listing both ways to buy. Nothing reaches your provider.
What we guarantee
- A run's budget is reserved before each provider call, so concurrent calls can't spend the same dollar. The method and raw results are public: per-run budget benchmark.
- Your provider key is read from each request and never written to disk, logged, or returned.
- Receipts hold costs and identifiers, not prompts or responses.
- Each payment adds credits at most once. A card purchase counts only after Stripe confirms it, and a USDC purchase only after it settles.
Experimental. These prices are a first offer and may change; credits you've already bought keep their terms.