Cost comparator
Put models side by side — on the cost that actually shows up on the bill.
Not just input/output. Model cached input, cache writes,reasoning tokens, and context tiers — then project it to a monthly spend at your request volume. Unknown rates stay unknown. Your selection is in the URL, so this comparison is a link you can share.
Tip: add 2–6 models. Try a quick pick:GPT-4o vs Claude Sonnet 4 ·GPT-4o mini vs Gemini 2.5 Flash
Tokens per request
Estimate tokens from a sample prompt
Buckets are disjoint: Input is fresh, uncached prompt tokens. Cached / cache-write / reasoning are billed separately.
Cost vs quality
cost/request now · quality = selected benchmark · up-and-left is better valueQuality: Epoch AI — “AI Benchmarking Hub” (epoch.ai), CC-BY. Third-party results, not a TokenTriage measurement.
Benchmarks: Epoch AI — “AI Benchmarking Hub” (epoch.ai), CC-BY. Third-party results, not a TokenTriage measurement. “Cost per passing result” = cost per request ÷ benchmark pass-rate — an illustrative value heuristic, not a promise about your workload.
The cheapest sticker price isn't the cheapest bill.
Cost per token is a proxy. TokenTriage measures cost per successful task on your real traffic — and only routes to the cheaper model once it's proven just as good, with evidence you can audit.