
Workload cost index
The cheapest route depends on the workload.
Every text model ranked on one fixed agent turn, so low input rates cannot hide expensive output. Same workload, same math, no markup.
100K
Input tokens
10K
Output tokens
Fast reads
Four useful price positions.
These picks are derived only from live catalog fields: price, context, and declared vision support. They are not editorial quality claims.
Full cost ledger
Ranked by the fixed turn.
Blended $/1M expresses the same 10:1 token mix as one normalized unit. Image-generation routes are excluded.
| Rank | Model route | Input | Output | Blended | Agent turn |
|---|---|---|---|---|---|
| 01 | DeepSeek V4 Flash DeepSeek / 1M | $0.09 | $0.18 | $0.098 | $0.011 |
| 02 | Muse Spark 1.2 Contributor Meta / 1M | $0.10 | $0.20 | $0.109 | $0.012 |
| 03 | Muse Spark 1.3 Contributor Meta / 1M | $0.10 | $0.20 | $0.109 | $0.012 |
| 04 | MiMo V2.5 Xiaomi / 1M | $0.11 | $0.28 | $0.125 | $0.014 |
| 05 | Gemini 2.5 Flash Lite Google / 1M | $0.10 | $0.40 | $0.127 | $0.014 |
| 06 | GPT-6 Luna OpenAI / 922K | $0.10 | $0.50 | $0.136 | $0.015 |
| 07 | MiMo V2.6 Flash Xiaomi / 1M | $0.14 | $0.28 | $0.153 | $0.017 |
| 08 | Qwen3.8 Flash Alibaba / 1M | $0.15 | $0.47 | $0.179 | $0.020 |
| 09 | GLM-5.3 Flash Zhipu / 1M | $0.15 | $0.50 | $0.182 | $0.020 |
| 10 | DeepSeek V4.1 Flash DeepSeek / 1M | $0.18 | $0.73 | $0.23 | $0.025 |
| 11 | Hy3 Tencent Hunyuan / 256K | $0.20 | $0.80 | $0.255 | $0.028 |
| 12 | DeepSeek V4 Flash Vision DeepSeek / 1M | $0.22 | $0.66 | $0.26 | $0.029 |
| 13 | Step 3.7 Flash StepFun / 256K | $0.20 | $1.15 | $0.286 | $0.032 |
| 14 | GPT-5.6 Luna OpenAI / 258K | $0.20 | $1.20 | $0.291 | $0.032 |
| 15 | Gemini 3.1 Flash Lite Google / 1M | $0.25 | $1.50 | $0.364 | $0.040 |
| 16 | LongCat 2.0 LongCat / 1M | $0.30 | $1.20 | $0.382 | $0.042 |
| 17 | MiniMax M2.7 MiniMax / 196K | $0.30 | $1.20 | $0.382 | $0.042 |
| 18 | MiniMax M3 MiniMax / 1M | $0.30 | $1.20 | $0.382 | $0.042 |
| 19 | Qwen3.7 Plus Alibaba / 1M | $0.32 | $1.28 | $0.407 | $0.045 |
| 20 | DeepSeek V4 Pro DeepSeek / 1M | $0.435 | $0.87 | $0.475 | $0.052 |
| 21 | MiMo V2.6 Pro Xiaomi / 1M | $0.435 | $0.87 | $0.475 | $0.052 |
| 22 | MiMo V2.5 Pro Xiaomi / 1M | $0.44 | $0.87 | $0.479 | $0.053 |
| 23 | Gemini 2.5 Flash Google / 1M | $0.30 | $2.50 | $0.50 | $0.055 |
| 24 | MiniMax M3 Highspeed MiniMax / 1M | $0.45 | $1.80 | $0.573 | $0.063 |
| 25 | Qwen3.6 Plus Alibaba / 1M | $0.50 | $3.00 | $0.727 | $0.080 |
| 26 | Gemini 3 Flash Google / 1M | $0.50 | $3.00 | $0.727 | $0.080 |
| 27 | MiniMax M2.7 Highspeed MiniMax / 196K | $0.60 | $2.40 | $0.764 | $0.084 |
| 28 | Doubao Seed 2.0 Code ByteDance / 200K | $0.67 | $3.36 | $0.915 | $0.10 |
| 29 | Doubao Seed 2.0 Pro ByteDance / 128K | $0.67 | $3.36 | $0.915 | $0.10 |
| 30 | Hy4 Preview Tencent Hunyuan / 1M | $0.834 | $2.501 | $0.986 | $0.11 |
| 31 | Gemini 3.6 Flash Google / 1M | $0.75 | $3.75 | $1.023 | $0.11 |
| 32 | Gemini 3.7 Flash Google / 1M | $0.75 | $3.75 | $1.023 | $0.11 |
| 33 | Gemini 3.8 Flash Google / 1M | $0.75 | $3.75 | $1.023 | $0.11 |
| 34 | Kimi K2.6 Moonshot / 256K | $0.95 | $4.00 | $1.227 | $0.14 |
| 35 | Kimi K2.7 Moonshot / 256K | $0.95 | $4.00 | $1.227 | $0.14 |
| 36 | Claude Haiku 4.5 Anthropic / 256K | $1.00 | $5.00 | $1.364 | $0.15 |
| 37 | GLM-5.1 Zhipu / 256K | $1.40 | $4.40 | $1.673 | $0.18 |
| 38 | GLM-5.2 Zhipu / 1M | $1.40 | $4.40 | $1.673 | $0.18 |
| 39 | GLM-5.3 Zhipu / 1M | $1.40 | $4.40 | $1.673 | $0.18 |
| 40 | Gemini 3.5 Flash Google / 1M | $1.50 | $9.00 | $2.182 | $0.24 |
| 41 | Qwen3.8 Max Alibaba / 1M | $2.00 | $6.00 | $2.364 | $0.26 |
| 42 | Grok 4.5 xAI / 500K | $2.00 | $6.00 | $2.364 | $0.26 |
| 43 | Grok 4.6 xAI / 500K | $2.00 | $6.00 | $2.364 | $0.26 |
| 44 | Grok 4.7 xAI / 500K | $2.00 | $6.00 | $2.364 | $0.26 |
| 45 | Claude Sonnet 5 Anthropic / 1M | $2.00 | $10.00 | $2.727 | $0.30 |
| 46 | GPT-6 Sol OpenAI / 922K | $2.00 | $10.00 | $2.727 | $0.30 |
| 47 | Gemini 3.1 Pro Google / 1M | $2.00 | $12.00 | $2.909 | $0.32 |
| 48 | GPT-5.6 Terra OpenAI / 258K | $2.00 | $12.00 | $2.909 | $0.32 |
| 49 | Qwen3.7 Max Alibaba / 1M | $2.50 | $7.50 | $2.955 | $0.33 |
| 50 | GPT-5.4 OpenAI / 1M | $2.50 | $15.00 | $3.636 | $0.40 |
| 51 | Claude Sonnet 4.6 Anthropic / 1M | $3.00 | $15.00 | $4.091 | $0.45 |
| 52 | Kimi K3 Moonshot / 1M | $3.00 | $15.00 | $4.091 | $0.45 |
| 53 | Claude Opus 4.6 Anthropic / 1M | $5.00 | $25.00 | $6.818 | $0.75 |
| 54 | Claude Opus 4.7 Anthropic / 1M | $5.00 | $25.00 | $6.818 | $0.75 |
| 55 | Claude Opus 4.8 Anthropic / 1M | $5.00 | $25.00 | $6.818 | $0.75 |
| 56 | Claude Opus 5 Anthropic / 1M | $5.00 | $25.00 | $6.818 | $0.75 |
| 57 | GPT-5.5 OpenAI / 256K | $5.00 | $30.00 | $7.273 | $0.80 |
| 58 | GPT-5.6 Sol OpenAI / 258K | $5.00 | $30.00 | $7.273 | $0.80 |
| 59 | Claude Fable 5 Anthropic / 1M | $10.00 | $50.00 | $13.636 | $1.50 |
| 60 | GPT-6 Astra OpenAI / 922K | $10.00 | $50.00 | $13.636 | $1.50 |
Single-rate ledgers
Read-heavy or write-heavy.
Cheapest input tokens
For large-file reads, retrieval, and context-heavy workloads.
Cheapest output tokens
For reasoning, code generation, and long-form responses.
Method / cost-index-v1
Comparable by construction.
One workload, every route
Each model gets the same 100K input and 10K output token profile. The ranking applies live catalog rates directly, rather than sorting on whichever headline number looks smallest.
Price is not capability
Image-generation routes are excluded because their billing is not comparable to a chat turn. Quality remains a separate decision: a cheap model that needs three attempts can cost more than a stronger route that succeeds once. Family rates live on the OpenAI pricing hub and the Anthropic pricing hub. Worked examples for two cheap catalog routes: Kimi K3 and Qwen3.8 Flash.
Before switching
Fair questions about cheap routes.
01What is the cheapest AI model API?+
On RouterPlex the cheapest model is DeepSeek V4 Flash at $0.09 per 1M input tokens and $0.18 per 1M output tokens — about $0.011 for a 100K-token-in, 10K-token-out agent turn. Prices are vendor list rates, so the same model costs the same whether you call it here or direct.
02Why rank by blended cost instead of input price?+
Ranking on input price alone favours models that charge little to read and a lot to write, which is the wrong way round for coding and agent work. This table ranks on a fixed 100K-in / 10K-out turn so input and output rates both count, and the same profile is applied to every model.
03How much cheaper is a budget model than a frontier model?+
Across this catalog the spread is about 139× on the same agent turn: DeepSeek V4 Flash costs roughly $0.011 where GPT-6 Astra costs about $1.50. Whether that trade is worth it depends on task difficulty — a cheap model that needs three attempts is not cheap.
04Is there a cheaper way to buy these models?+
Not meaningfully. RouterPlex bills vendor list prices with no markup and no top-up fee, so per-token cost matches going direct to each vendor. What you save is the overhead of separate accounts, keys and minimum balances per provider.
05Do cheap models have rate limits or quality caveats?+
RouterPlex applies no additional RPM or TPM cap beyond the upstream vendor's. The real caveat is capability: budget models are generally weaker at long agentic tool loops, so compare quality on the benchmarks pages before moving production traffic onto one.
Route on your own math
One key reaches every model in the ledger.
Top up from $5 and pay the exact vendor token rate.