GitHubBook a demoStart routing

The honest LLM gateway

Cut your LLM bill.
Keep the proof.

Point your apps at one base URL to see where every dollar goes. Test cheaper models in shadow, and route only after quality holds — so every saving is proven, estimated, or honestly labeled can’t-prove-it-yet. Never a number we can’t back up.

git clone https://github.com/inferops/tokentriage && cd tokentriage
CGO_ENABLED=0 go build -o dist/tokentriage ./cmd/tokentriage
./dist/tokentriage run --anthropic

Build from source · data plane on :8787

  • Source-available (FSL)
  • Self-host in one command
  • No signup
TokenTriage routing illustrationRequest cards are scanned, tagged with the policies they matched, and switched to one of three model tiers, with a live routing log recording each decision.gpt-5 minifastestsonnet 5balancedopus 4.8top qualityyourssupport replylow costfast response
live routing log
analysistop quality→ sonnet 5failover
greetingexact match→ cacheno model call
code reviewlong context→ sonnet 5
Illustrative routing simulation — no real traffic.
2,858models in the price catalog

One base URL for everything you already run.

  • OpenAI
  • Anthropic
  • Azure OpenAI
  • AWS Bedrock
  • Google Gemini / Vertex
  • Claude on Vertex
  • Mistral
  • DeepSeek
  • xAI (Grok)
  • Groq
  • Fireworks
  • OpenRouter
  • Perplexity
  • Cohere
  • Together AI
  • ElevenLabs (TTS)
  • Deepgram (STT)
  • Local (vLLM · Ollama)

80+ providers, one gateway — every request priced and attributed from the same catalog.

Head-to-head · measured on bare metal · pinned versions

62× less overhead than LiteLLM. Measured, not marketed.

Same bare-metal run — AWS z1d.metal, an equal 4-core budget, latest stable images, steel-manned configs — and TokenTriage was computing and logging cost for every request while LiteLLM ran its logging-off latency posture. Every number below is reproducible from the open harness.

62×less added latency than LiteLLMmedian gateway overhead, same run — and 2.6× less than Bifrost

Under load, it isn't close.

Offered load swept 100 → 4,000 req/s on the same 4 cores. LiteLLM saturates at ~670 req/s and its p99 collapses past 13 seconds. TokenTriage takes all 4,000 at 1.20 ms.

p99 latency versus offered load: TokenTriage, Bifrost, and LiteLLM on the same 4-core budgetA log-scale line chart. TokenTriage and Bifrost hug the floor across the whole sweep — TokenTriage ends at 1.20 milliseconds p99 at 4,000 requests per second, Bifrost at 3.54 milliseconds, with one disclosed TokenTriage transient of 63 milliseconds at 2,000. LiteLLM climbs from 6.5 milliseconds at 100 requests per second to 42 at 500, then saturates at about 670 achieved requests per second and its p99 collapses to between 13 and 17 seconds.1 ms10 ms100 ms1 s10 s1005001k2k4koffered load (requests / second)saturated: ~670 req/s achievedone transient: 63 msLiteLLM16.8 s p99Bifrost · 3.54 msTokenTriage · 1.20 ms
p99 latency vs. offered load (both axes log). Lower is better. Same run, same 4-core budget, pinned images: LiteLLM v1.94.0 · Bifrost v1.6.6.

The current optimized build goes further (TokenTriage-only follow-up run, same hardware class — kept separate; we don't splice runs into the head-to-head)

  • 8,000+ req/s sustained · 100% success
  • 1.45 ms p99 at 8,000 req/s
  • ~10,200 req/s at peak, still 100%

Cost accuracy: a three-way tie. Auditability: no contest.

We pre-committed to publish this axis win or lose: TokenTriage, LiteLLM, and Bifrost all matched an independent provider-doc reference to the nanodollar. The difference is what happens after: TokenTriage stamps every request with its pricing snapshot and decision provenance, so each decimal survives an audit.

Reported cost, all three gateways = ground truth

  • 0% cached$0.006400000 3/3 exact
  • 25% cached$0.006112000 3/3 exact
  • 50% cached$0.005824000 3/3 exact

AWS z1d.metal · 48 vCPU bare metal · equal 4-core gateway budgets · disjoint-core pinned · deterministic mock upstream · 95% bootstrap CIs · LiteLLM v1.94.0 · Bifrost v1.6.6 · pinned image digests

See the full head-to-head & reproduce it yourself

The problem

LLM spend is exploding — and opaque. Cutting it blind risks quality.

If more than one of these is true, there is almost certainly a number worth finding.

What it does

Route · Attribute · Prove · Govern · Observe

One gateway your apps point at. Each capability names the outcome first — the mechanism is one layer down for the engineers.

Route

Send each request to the right model

Point each request at the cheapest tier that can still do the job — hard prompts to a frontier model, easy ones to a cheap one. Routing stays off until you turn it on.

11 routing strategies · typed policy language · off by default

See how routing decides
Demo dataTokenTriage routing cockpit “Live rules” table in demo mode: each rule (pin, override, default, simple-to-cheap) with its hit count, realized cost, and would-route savings — EST flags on estimated figures, unpriced rows counted. Watermarked DEMO DATA.
Real dashboard, demo mode — each routing rule with its realized cost and the savings it would capture, EST-flagged where estimated.

Attribute

See where every dollar goes

Every request is costed and attributed by model, token type, trace, agent, and MCP tool — all traceable to one decision log. Nothing is guessed; unpriced spend is counted, never shown as $0.00.

One JSONL decision log · source@date#sha7 pricing · UNPRICED ≠ $0.00

See where dollars go
Demo dataTokenTriage cost overview in demo mode: total spend of priced rows (EST), request, model, and tenant counts, a mixed-snapshot notice, and a spend-over-time chart — unpriced rows excluded from every total. Watermarked DEMO DATA.
Real dashboard, demo mode — every dollar attributed, with unpriced rows excluded from totals and estimated figures flagged EST.

Prove

Only cut costs after the cheaper path is proven

Test a cheaper model in shadow against your real traffic and grade its quality before anything changes. Every saving is labeled PROVEN, ESTIMATED, or honestly can’t-prove-it-yet.

Verdicts: PROVEN / ESTIMATED / UNPROVABLE · shadow-judged, never guessed

Read a real verdict
Demo dataTokenTriage routing cockpit trust-ladder in demo mode: Shadow and Oracle verdicts read CANNOT_MEASURE and Outcome reads ON_FRONTIER, with a note that served-model labels can't show routing skill — enable shadow duplication for a counterfactual verdict. Watermarked DEMO DATA.
Real dashboard, demo mode — it refuses to claim a routing win it cannot measure yet.

Govern

Cap spend before it happens

Mint virtual keys (hashed, shown once) and enforce per-tenant budgets and rate limits across org, team, user, and key. When the limiter itself dies it fails open and discloses it — never a fake 429.

6 limit types · virtual keys (hashed, shown once) · fail-open with disclosure

Set up a virtual key
Demo dataTokenTriage budget “prod-monthly” in demo mode: a spend bar at 42.1% of an alert-only budget, a burn-rate exhaustion estimate, and a callout that an alert-only budget tracks burn but never refuses a request — unpriced rows excluded from the total.
Real dashboard, demo mode — an alert-only budget tracks burn and discloses that it never stops spending.

Observe

Watch cost and quality in real time

A dashboard that reads the same decision log the ledger writes — spend, savings, and health, live. It sits entirely off the request hot path, so a dead dashboard can never slow your traffic.

OTLP tokentriage.cost.* · dashboard off the hot path

Watch cost live
Demo dataTokenTriage spend panel in demo mode: spend to date, a projected period-end (EST, linear from the last 24 hours), and savings vs. baseline reading “No verdict yet — savings appear once a shadow comparison is graded.” Watermarked DEMO DATA.
Real dashboard, demo mode — live spend and a projection, with savings withheld until a shadow comparison is graded.

How it works

Climb the trust ladder — and never claim a win you can’t show

You earn enforcement one rung at a time: observe every dollar, prove the cheaper path in shadow, then route. The same decision can only ever be proven, estimated, or honestly refused.

  1. 1Ledger

    See every dollar

    Cost and log every request. Zero request mutation.

  2. 2Shadow

    Prove it is safe

    Compute the cheaper counterfactual and judge its quality — without touching a response.

  3. 3Route

    Enforce, once earned

    Switch to the cheaper tier only after the evidence clears the bar.

One decision, three honest answers

  • Proven

    The shadow judge graded the cheaper tier on the same query and it held. The saving is real.

  • Estimated

    Only the counterfactual cost (would_usd) is known, not a quality judgement yet. The number wears an EST flag.

  • Unprovable

    With served-model labels alone, routing can at best sit on the frontier — never a beat. An honest refusal, never a number.

Demo dataTokenTriage routing cockpit in demo mode: trust-ladder verdicts reading CANNOT_MEASURE, and a gated 'Promote to route' with the evidence-acknowledgement it requires — watermarked DEMO DATA.
The real routing cockpit, demo mode: promotion to Route stays gated until the evidence is acknowledged — the ladder won’t let you claim a win it can’t show.

See how it works

The trust dashboard

A dashboard where every number survives an audit

Spend to date, projected period-end, and savings vs. baseline — each stamped with its pricing snapshot, EST flags where estimated, and UNPRICED rows counted, never hidden.

localhost:9090/uiDemo data
TokenTriage overview, zoomed to the cost panel: spend to date $63.38, projected period-end, and savings vs. baseline, marked DEMO DATA.TokenTriage overview, zoomed to the cost panel: spend to date $63.38, projected period-end, and savings vs. baseline, marked DEMO DATA.
The real TokenTriage dashboard in demo mode, cropped to the cost-attribution panel. The DEMO DATA watermark and sample figures come from a synthetic ledger — the product ships this so you can explore before wiring real traffic.

Open at the core

Read the code. Audit it yourself.

The core is source-available under the Functional Source License and the client SDKs are Apache-2.0. Every honesty invariant on this page is a constraint in the source — so you don’t have to take our word for any of it.

Core
FSL-1.1-ALv2 — converts to Apache-2.0 two years after each release
Client SDKs
Apache-2.0 (Python · TypeScript)
Deploy
One ≈10.7 MB static binary · self-host · air-gap safe

Pricing

Free to self-host. Honest gates only.

See where your LLM spend goes

Cut your LLM bill only after you prove the cheaper path is just as good.

Building in the open

Get launch updates

TokenTriage is being built in the open. Drop your email for occasional, no-spam updates — new releases, fresh benchmarks, and the road to 1.0. Star the repo for the code; join the list for the story.

Double opt-in. Unsubscribe anytime. We never share your address.