Inferrail

For agents and the people who run them

Inferrail HostedOne budget per run, across your whole fleet.

Give each AI run a dollar ceiling in a request header. Every call in that run shares it, including parallel calls from different agents and machines, and a call that would push the run past it is refused before it reaches OpenAI or Anthropic. You bring your own provider key with each request; we never store it.

Experimental · 2,000 governed runs a month free · then $0.001 per run · no subscription

Not open yet. This page describes the terms Inferrail Hosted will launch with. The hosted service isn't accepting workspaces yet; until it is, the open-source gateway does the same per-run enforcement on your own machine.

Price

Unit

A governed run: a distinct work id (X-Inferrail-Attribute-Work-Id) in a calendar month (UTC). Calls inside a run are not charged separately.

Billed when

At least one call in the run is answered by your provider. A run whose calls all fail, or are all refused by its own budget, costs nothing.

Free

2,000 governed runs per workspace per month.

After that

Prepaid credits at $0.001 per run. People pay by card: $10 for 10,000 runs. Agents pay in USDC over x402: $1 for 1,000 runs, with no account.

No surprises

Credits are prepaid. No subscription, no auto top-up, no overage bill. When the allowance and credits run out, calls are refused before reaching your provider, with a link to buy more.

Model spend

Billed by your provider on your own key. Inferrail doesn't resell model access or mark it up.

Self-hosted

The open-source gateway is free and does everything on one machine: pip install inferrail. Hosted is for when many agents and machines must share one budget ledger and you'd rather not run it.

Start

Past your allowance with no credits left, a call gets HTTP 402 allowance_exhausted, listing both ways to buy. Nothing reaches your provider.

What we guarantee

Experimental. These prices are a first offer and may change; credits you've already bought keep their terms.