GitHubBook a demoStart routing
MI

Minimax M3

MinimaxChat1M context· released 2026-06-12Floating alias

minimax/MiniMax-M3

Prices as of 2026-07-14 · rev 179affb· unknown is shown as unknown, never $0.00

Data confidence1/1 exactbenchmarks measured on this exact model · prices 2026-07-14
Cost / request$0.00091K in · 500 out
Input / M$0.30
Output / M$1.20
Cache read / M$0.06
Context1M1,000,000 tokens
Max output128K128,000 tokens

Capability fingerprint

independent benchmarks, grouped by domain · higher is better · not our measurement

Reasoning

SimpleBencheveryday reasoning (unsaturated)
35%53rd pctile
exact modelSimpleBench Leaderboard · MiniMax-M3

peer median  ·  percentile ranks this model among distinct comparable models (deduped — one entry per model, not per provider) scored on that benchmark  ·  exact model / family proxy per row.

Source: Epoch AI — “AI Benchmarking Hub” (epoch.ai), CC-BY (Epoch AI). Some rows carry an upstream leaderboard (shown per row); Aider & Terminal-Bench are additionally Apache-2.0. Each benchmark measures a different thing on a different scale — rows are not comparable across domains, and a score is not a guarantee of quality on your task.

Capabilities Index

Epoch AI's cross-model capability re-fit · one comparable axis · not our measurement
147.0ECI
72nd percentile of 150 modelsexact model
84.0111.4138.7166.1Epoch Capabilities Index (relative re-fit scale)peer median 139.7: 147.0 (95% CI 141.9–149.6)147.0

The dot is Minimax M3; the whisker is its 95% CI; the firm tick is the peer median; faint ticks are every other scored model. Unlike a raw benchmark, the ECI axis is comparable across models — but it's a relative re-fit scale, so a model's exact position can shift when Epoch recomputes it.

Source: Epoch AI — Capabilities Index & Notable-Models data (epoch.ai), CC-BY (Epoch AI). The ECI is Epoch's index, not TokenTriage's; per-fact confidence labels are Epoch's own.

Human preference (LMArena)

blind pairwise human votes · Elo · style-controlled · not a correctness score
exact modelmeasured on minimax-m3statistically tied with 16 models overall
1065123714081580Arena score (Elo · higher = more often preferred)Overallpeer median 1357Overall: 1445 (95% CI 1439–1450)1445#66 · 30K votesCodingpeer median 1400Coding: 1498 (95% CI 1491–1506)1498#57 · 8.7K votesHard promptspeer median 1370Hard prompts: 1466 (95% CI 1460–1472)1466#63 · 20K votes

Each dot is Minimax M3's Arena score in that category; the whisker is the 95% CI; the firm tick is the peer median across distinct models; faint ticks are the field. This is human preference — which answer people pick in a blind A/B — not a correctness or capability score, and adjacent ranks routinely overlap within their intervals.

Source: LMArena (LMArena Leaderboard (lmarena.ai) — lmarena-ai/leaderboard-dataset on Hugging Face, CC BY 4.0.). Ratings are the style-controlled estimates (length/formatting confound removed), as of 2026-07-27.

Also worth considering

cheaper models that score about as well — pick the benchmark that matters to you

1 cheaper model scores within 3 points of Minimax M3 on SimpleBench.

ModelSimpleBenchCost / reqSavings
DEDeepseek V4 Pro53%+18$0.00087−3%

Deduped to one entry per model (cheapest offering), priced at the reference 1K-in / 500-out request. Filtered only on the benchmark score — a model with an unknown capability is never silently excluded. A lower price is only cheaper if quality holds on your task — that's what routing proves.

Pricing

Input · fresh prompt tokens$0.30 / M
Output · generated tokens$1.20 / M
Cache read · cached-input hit$0.06 / M
Cache write · cache creationunknown
Reasoning · thinking tokensunknown

Context-length tiers

Base (≤ 512K)$0.30 / $1.20 per M
Above 512K$0.60 / $2.40 per M

When total input exceeds the threshold, the tier rate applies to the whole request.

$0.65$0.49$0.32$0.16$00250K500K750K1MTotal input tokensCost / requestthe whole request reprices at 512K tokens (+100%)whole request reprices +100% at 512Krates as of 2026-07-14 · rev 179affb

Cross the 512K line by a single token and the whole request reprices — every token, not just those above the line. Assumes 500 output tokens (the cliff's position depends on input alone). Snapshot rates; providers do change tier structure.

Reasoning tokens: no separate rate is published. Providers typically bill thinking tokens at the output rate ($1.20 / M) — which is what the estimator assumes, so an extended-thinking workload isn't silently undercounted.

Model info

Provider
Minimax
Catalog key
minimax/minimax/MiniMax-M3
Identifier
Floating alias (may re-point over time)
Type
Chat
Max input
1,000,000 tokens
Max output
128,000 tokens
Released
2026-06-12
Knowledge cutoff
unknown
Weights
Open weights
Price snapshot
2026-07-14 · rev 179affb

Pricing from TokenTriage's resolved price artifact (TokenTriage resolved price artifact (internal/pricing/data/prices.json)); capability + model metadata is third-party (LiteLLM model_prices_and_context_window.json, models.dev). Unknown means unstated, not absent.

Capabilities

4 documented as supported

Input & output

Vision
?PDF input
?Audio input
?Audio output

API behaviour

?Streaming
?Structured output
Prompt caching
Reasoning

Agents & tools

Function calling
?Parallel tool calls
?Web search
?Computer use

supported · not supported ·? not documented · sources disagree. A models.dev tag means the LiteLLM catalog was silent and an independent source supplied the value (lower confidence); ? is never silently turned into a confident ✕.

Open weights

verified Hugging Face repo · facts from the model's safetensors index

Model

Parameters
427.0B
Architecture
MiniMaxM3SparseForConditionalGeneration
Native precision
BF16

Weights footprint

Native size
854 GB

Memory to hold the weights — not VRAM. Real serving needs extra headroom for activations and the KV cache, which grows with context length and batch size.

License & source

License
otherCommercial use — conditional

Source: Hugging Face Hub — facts read from the model repo as of 2026-07-30. Each fact carries the model's own repo license (shown above); TokenTriage neither hosts nor relicenses the weights.

Estimate a request

disjoint token buckets · input = fresh (uncached) tokens
Preset
Cost / request
Projected / month

How output length drives the bill

at a 1K-token prompt · cost scales linearly with output
$0.04$0.03$0.02$0.01$0032K64K96K128KOutput tokensCost / requestmax outputoutput spend = input spend at ~250 tokensoutput spend overtakes input at ~250

Output tokens overtake input spend at ~250 tokens for this model — past that, generation is your bill, whatever the headline input price says. Computed by the same cost engine as the estimate above.

Route to Minimax M3

OpenAI-compatible · one base URL
curl https://YOUR-TOKENTRIAGE-HOST/v1/chat/completions \
  -H "Authorization: Bearer $TOKENTRIAGE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "minimax/minimax/MiniMax-M3",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Point your existing OpenAI client at your TokenTriage host and pass the catalog key as the model. TokenTriage attributes the real cost per request and can route to a cheaper model once it's proven just as good on your traffic.

Data & API

machine-readable · rev 179affb

Pull this model's pricing, capabilities, benchmarks and cross-source provenance programmatically — every field carries the snapshot rev, and unknown is null, never 0. Free to reuse with attribution: benchmarks & Capabilities Index © Epoch AI (CC-BY (Epoch AI)); human-preference Elo © LMArena (CC-BY-4.0); capability & price cross-check via models.dev v2 (MIT) · Portkey (MIT) · TrueFoundry (MIT); open-weights facts via Hugging Face (per-repo license).

A price is a guess until it meets your traffic.

This page tells you what Minimax M3 charges per token. TokenTriage tells you what it costs on your real requests — and cuts the bill only after proving a cheaper model is just as good.