GitHubBook a demoStart routing
OP

GPT Realtime 2.1

OpenAIChat128K context· released 2026-07-09Floating alias

gpt-realtime-2.1

Prices as of 2026-07-14 · rev 179affb· unknown is shown as unknown, never $0.00

Cost / request$0.0161K in · 500 out
Input / M$4.00
Output / M$24.00
Cache read / M$0.40
Context128K128,000 tokens
Max output32K32,000 tokens

Pricing

Input · fresh prompt tokens$4.00 / M✓ 2 sources agree
Output · generated tokens$24.00 / M✓ 2 sources agree
Cache read · cached-input hit$0.40 / M
Cache write · cache creationunknown
Reasoning · thinking tokensunknown

Model info

Provider
OpenAI
Catalog key
openai/gpt-realtime-2.1
Identifier
Floating alias (may re-point over time)
Type
Chat
Max input
128,000 tokens
Max output
32,000 tokens
Released
2026-07-09
Knowledge cutoff
2024-09-30
Weights
Proprietary
Price snapshot
2026-07-14 · rev 179affb

Pricing from TokenTriage's resolved price artifact (TokenTriage resolved price artifact (internal/pricing/data/prices.json)); capability + model metadata is third-party (LiteLLM model_prices_and_context_window.json, models.dev). Unknown means unstated, not absent.

Capabilities

4 documented as supported · 1 source conflict

Input & output

Visionmodels.dev
?PDF input
Audio input
Audio output

API behaviour

?Streaming
Structured outputmodels.dev
?Prompt caching
Reasoningmodels.dev

Agents & tools

Function callingconflict
Parallel tool calls
?Web search
?Computer use

supported · not supported ·? not documented · sources disagree. A models.dev tag means the LiteLLM catalog was silent and an independent source supplied the value (lower confidence); ? is never silently turned into a confident ✕.

Estimate a request

disjoint token buckets · input = fresh (uncached) tokens
Preset
Cost / request
Projected / month

How output length drives the bill

at a 1K-token prompt · cost scales linearly with output
$0.82$0.61$0.41$0.2$008K16K24K32KOutput tokensCost / requestmax outputoutput spend = input spend at ~167 tokensoutput spend overtakes input at ~167

Output tokens overtake input spend at ~167 tokens for this model — past that, generation is your bill, whatever the headline input price says. Computed by the same cost engine as the estimate above.

Route to GPT Realtime 2.1

OpenAI-compatible · one base URL
curl https://YOUR-TOKENTRIAGE-HOST/v1/chat/completions \
  -H "Authorization: Bearer $TOKENTRIAGE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-realtime-2.1",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Point your existing OpenAI client at your TokenTriage host and pass the catalog key as the model. TokenTriage attributes the real cost per request and can route to a cheaper model once it's proven just as good on your traffic.

Data & API

machine-readable · rev 179affb

Pull this model's pricing, capabilities, benchmarks and cross-source provenance programmatically — every field carries the snapshot rev, and unknown is null, never 0. Free to reuse with attribution: benchmarks & Capabilities Index © Epoch AI (CC-BY (Epoch AI)); human-preference Elo © LMArena (CC-BY-4.0); capability & price cross-check via models.dev v2 (MIT) · Portkey (MIT) · TrueFoundry (MIT); open-weights facts via Hugging Face (per-repo license).

A price is a guess until it meets your traffic.

This page tells you what GPT Realtime 2.1 charges per token. TokenTriage tells you what it costs on your real requests — and cuts the bill only after proving a cheaper model is just as good.