inferrail

Inferrail — Product

This file is the authoritative source for exact current scope.

Developer Preview. The scope below is fully implemented and tested, but nothing is stable yet — CLI flags, inferrail.yaml’s shape, and receipt/telemetry JSON fields may change without notice before v1.0. The package version lives in pyproject.toml (and on PyPI); it is not restated here so it cannot drift.

What it is

Inferrail is an open inference control plane: infrastructure that sits between an application and the model providers it calls, so operational decisions (which provider, which model, retry/fallback behavior, what happened and why) live in configuration and telemetry rather than scattered across application code.

v0.1 is the first, narrow slice of that: an OpenAI-compatible HTTP gateway that takes a chat completion request, routes it to one explicitly configured provider/model via a static policy, executes it, and returns a compatible response plus a structured local telemetry record of what happened — plus, as of this slice, a privacy-preserving economic receipt: what that execution cost, computed from measured usage and verified pricing, tied to business context the caller attaches. The long-term thesis this is the first step of: measure → attribute → connect to outcome → govern → optimize.

v0.2.0 adds a second, separate product on top of that same substrate: AP invoice-exception recovery — see the next section. It is a bounded decision-and-execution engine for one specific operational question (retry vs. human review), not a general invoice-processing product, and it is fully isolated from the gateway and from Work Economics/Economic Authority below (own module, own storage, no shared code path).

AP invoice-exception recovery (v0.2.0)

For teams operating an invoice-extraction workflow: decide whether one eligible extraction exception gets one permitted machine retry or your established human-review path, execute the retry through a supported integration, and record the resulting cost and outcome.

Zero customer adoption or savings claims are made about this capability — see the capability doc’s “Pricing and performance assumptions” section, which labels every dollar figure as an assumption, not a validated result.

Who it’s for

Developers and small teams running LLM-backed applications who want:

The problem being solved right now

Applications that call LLM providers directly hard-code operational decisions (which provider, which model, how to handle a 429) into business logic. Inferrail v0.1 moves that decision to a deterministic, inspectable config file and gives you a telemetry record for every request — the prerequisite for anything smarter later (see “Long-term direction”).

Current scope: what works today

Cost and receipts

Budgets and enforcement

See docs/adr/0015-budget-enforcement.md for the full design; summary here:

Local control API and --app-mode

See docs/adr/0016-local-control-api.md for the full design; summary here — this is a local, single-process API (not the hosted, cross- fleet “control plane” docs/adr/0004 anticipates):

Dashboard (v0.4.0) — all six screens built

A static, local web SPA in app/, served by the same process as the local control API when --app-mode is on — see docs/adr/0017-dashboard-in-app-directory.md.

Usage/presence beacon (opt-out)

A real feature, not a placeholder (src/inferrail/usage_ping/, docs/adr/0020-quickstart-both-sdks-and-payload-free-verification.md, docs/privacy/usage-ping.md). On by default (a deliberate reversal of the original opt-in design — see ADR-0020, which supersedes docs/adr/0019-opt-in-usage-ping.md’s “default off” only), but still inert with no usage_ping.endpoint configured in inferrail.yaml regardless of the toggle — Inferrail ships with no built-in default endpoint, so an install is never able to send anything until an operator explicitly configures one. Fires for every inferrail serve invocation now, not just --app-mode (quickstart and plain config-based deployments included). Four lifecycle events only: install (once ever), serve_start (every process start), first_receipt (once ever), heartbeat (at most once per 24 hours while serving). Never a prompt, response, model name, cost, work_id, project name, or anything about actual traffic. Sending never blocks, slows, or can fail the gateway — it happens on a best-effort background thread with a short timeout, and any failure (offline, unreachable, timeout) is swallowed.

Turned off with inferrail telemetry disable, the Settings screen’s toggle, INFERRAIL_TELEMETRY=0, serve --no-telemetry, or DO_NOT_TRACK=1 — and automatically under common CI environment variables and this project’s own test suite, no configuration needed.

Two ways to verify what would be sent without trusting this documentation: inferrail telemetry preview (prints the exact payload for every event, from this install’s real id/OS/version, without sending anything) and the Settings screen’s own link to docs/privacy/usage-ping.md. inferrail telemetry status|enable|disable work standalone, without --app-mode or even an inferrail.yaml.

A reference collector (hosted/usage_ping/) is built — own process, own SQLite storage (an installs table tracking activation via reached_first_receipt_at, plus an append-only events log), zero dependency on the inferrail package, never logs or persists the connecting IP address, admin-gated aggregate /stats (disabled entirely unless an admin token is set). scripts/owner_stats.py reads PyPI download counts (always, no telemetry needed) plus, if pointed at a deployed collector’s database with --db, real install/activation/ active-user numbers. Deploying a collector instance and configuring usage_ping.endpoint to point at it is a human action — see PROGRESS.md’s “HUMAN ACTION NEEDED”.

Diagnostics: inferrail pricing update and inferrail doctor

Explicit non-goals / not yet supported

Not a hidden limitation — these are the honest edges of v0.1:

Verifying privacy claims yourself

The claim being checked: Inferrail’s receipt and telemetry paths do not copy prompt or response bodies into the records they write. Three ways to check it, each with a different scope:

  1. inferrail verify-payload-free lists every InferenceReceipt field and checks that none is named for message content. It inspects the schema only. attributes is a free-form dict[str, str] stored as the caller sends it, so a name check cannot prove that stored values are free of sensitive text.
  2. The canary tests send marker strings through the real gateway paths (success, provider error, streaming, tool calls, both /v1/chat/completions and /v1/messages) and assert the markers never reach a receipt or telemetry event: tests/unit/test_gateway_receipts.py, tests/unit/test_gateway.py, tests/unit/test_gateway_anthropic.py. The code those tests exercise: request handlers (src/inferrail/gateway/routes.py), execution engines (gateway/execution.py, gateway/anthropic_execution.py), provider adapters (providers/openai.py, providers/anthropic.py), the receipt builder (receipts/builder.py), and the sinks (receipts/sinks.py, receipts/sqlite_store.py), all under src/inferrail/.
  3. Against your own running gateway:

In inferrail.yaml, set:

telemetry:
  sink: jsonl
  path: inferrail-telemetry.jsonl

Restart inferrail serve, then:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "default", "messages": [{"role": "user", "content": "MARKER-1234-do-not-persist-me"}]}'

grep -c "MARKER-1234" inferrail-telemetry.jsonl   # 0, every time
cat inferrail-telemetry.jsonl                     # latency, tokens, status — no message content

grep -c "MARKER-1234" inferrail-receipts.jsonl    # 0, every time (receipts are on by default)
cat inferrail-receipts.jsonl                      # tokens, cost, pricing — no message content

This checks only what Inferrail itself writes to disk for that request. It does not cover values you put in attribution headers, your own application or proxy logs, or the provider: your provider still receives the real prompt. Inferrail is a pass-through gateway to it, not a privacy boundary against it. None of these checks is a security audit.

Hosted capabilities (beyond the self-hosted data plane)

Everything above is the self-hosted gateway: zero dependency on any Inferrail-operated service (see “Who it’s for” and docs/adr/0004).

Hosted cost gateway trial (hosted/cost_gateway/): a no-account trial at tryinferrail.com/try/ that issues a short-lived, isolated tenant with zero-key demo traffic, and optionally proxies real traffic through a visitor’s own OpenAI or Anthropic key. With a real key, the hosted process holds that key in memory and handles the traffic; the trial expires within 4 hours of the key being added (24 hours at most in demo mode). Preview status. Full contract and key-handling threat model: hosted/cost_gateway/README.md. Inferrail helps companies measure, attribute, and eventually govern the economics of work performed by AI agents. The gateway and receipts above are today’s working measurement layer. Separately, Inferrail also operates two hosted, paid capabilities that extend that foundation toward machine buyers. Both are experimental and Base Sepolia testnet only; neither controls external wallets, providers, or network spending:

Long-term direction

The progression Inferrail is built to support, in order, is:

observable → controllable → measurable → comparable → optimizable → increasingly intelligent

Each step needs the one before it to be trustworthy first. v0.1 delivered the first step (observable: a real telemetry record for every request) and the scaffolding for the second (controllable: explicit routing config, provider abstraction). This slice takes the first real step into measurable: a per-request economic receipt (verified pricing, Decimal cost, business attribution) and a local report to read it back — still a single-process, single-machine capability, not fleet-wide history. Later phases — comparing providers/models on cost and quality across many requests, recommending routing policies, fleet-wide analytics — depend on operational history accumulating across many requests/deployments, which is naturally a hosted capability once a user wants it to span more than one machine or process. See ARCHITECTURE.md for how the OSS data plane and a future hosted control plane are meant to stay decoupled.

Non-goals (for the project generally, not just this session)

Inferrail is not attempting to become: