GitHubBook a demoStart routing

← All models

Cost comparator

Put models side by side — on the cost that actually shows up on the bill.

Not just input/output. Model cached input, cache writes,reasoning tokens, and context tiers — then project it to a monthly spend at your request volume. Unknown rates stay unknown. Your selection is in the URL, so this comparison is a link you can share.

Prices: TokenTriage price snapshot · updated 2026-07-14rev 179affbMath mirrors TokenTriage's own cost engine

Tip: add 2–6 models. Try a quick pick:GPT-4o vs Claude Sonnet 4 ·GPT-4o mini vs Gemini 2.5 Flash

Tokens per request

Estimate tokens from a sample prompt
Rough estimate (~4 characters per token). For an exact count, use the provider's tokenizer.

Buckets are disjoint: Input is fresh, uncached prompt tokens. Cached / cache-write / reasoning are billed separately.

Add at least one model above to see the cost comparison. Or start from a quick pick.

The cheapest sticker price isn't the cheapest bill.

Cost per token is a proxy. TokenTriage measures cost per successful task on your real traffic — and only routes to the cheaper model once it's proven just as good, with evidence you can audit.