Cheapest AI models.
Ranked on what a real turn costs.

Every text model on RouterPlex, cheapest first. Most “cheapest model” tables sort on input price, which flatters models that charge little to read and a lot to write. This one ranks on a fixed 100K-in / 10K-out agent turn, so input and output rates both count.

Cheapest overall
DeepSeek V4 Flash

$0.011 per agent turn — the lowest blended cost in the catalog.

Cheapest with vision
Step 3.7 Flash

Accepts image input at $0.20 per 1M input tokens.

Largest context window
Kimi K3

1M context at $0.45 per agent turn.

#ModelProviderInput / 1MOutput / 1MBlended / 1MPer agent turn
1DeepSeek V4 Flash1MDeepSeek$0.090$0.18$0.098$0.011
2MiMo V2.51MXiaomi$0.11$0.28$0.13$0.014
3Hy3256KTencent Hunyuan$0.20$0.80$0.25$0.028
4Step 3.7 Flash256KStepFun$0.20$1.15$0.29$0.032
5LongCat 2.01MLongCat$0.30$1.20$0.38$0.042
6MiniMax M2.7196KMiniMax$0.30$1.20$0.38$0.042
7MiniMax M31MMiniMax$0.30$1.20$0.38$0.042
8Qwen3.7 Plus1MAlibaba$0.32$1.28$0.41$0.045
9DeepSeek V4 Pro1MDeepSeek$0.43$0.87$0.47$0.052
10MiMo V2.5 Pro1MXiaomi$0.44$0.87$0.48$0.053
11MiniMax M3 Highspeed1MMiniMax$0.45$1.80$0.57$0.063
12Qwen3.6 Plus1MAlibaba$0.50$3.00$0.73$0.080
13MiniMax M2.7 Highspeed196KMiniMax$0.60$2.40$0.76$0.084
14Doubao Seed 2.0 Code200KByteDance$0.67$3.36$0.91$0.10
15Doubao Seed 2.0 Pro128KByteDance$0.67$3.36$0.91$0.10
16Kimi K2.6256KMoonshot$0.95$4.00$1.23$0.14
17Kimi K2.7256KMoonshot$0.95$4.00$1.23$0.14
18Claude Haiku 4.5256KAnthropic$1.00$5.00$1.36$0.15
19GPT-5.6 Luna258KOpenAI$1.00$6.00$1.45$0.16
20GLM-5.1256KZhipu$1.40$4.40$1.67$0.18
21GLM-5.21MZhipu$1.40$4.40$1.67$0.18
22Gemini 3.5 Flash1MGoogle$1.50$9.00$2.18$0.24
23Qwen3.8 Max1MAlibaba$2.00$6.00$2.36$0.26
24Grok 4.5500KxAI$2.00$6.00$2.36$0.26
25Grok 4.6500KxAI$2.00$6.00$2.36$0.26
26Claude Sonnet 51MAnthropic$2.00$10.00$2.73$0.30
27Gemini 3.1 Pro1MGoogle$2.00$12.00$2.91$0.32
28Qwen3.7 Max1MAlibaba$2.50$7.50$2.95$0.33
29GPT-5.41MOpenAI$2.50$15.00$3.64$0.40
30GPT-5.6 Terra258KOpenAI$2.50$15.00$3.64$0.40
31Claude Sonnet 4.61MAnthropic$3.00$15.00$4.09$0.45
32Kimi K31MMoonshot$3.00$15.00$4.09$0.45
33Claude Opus 4.61MAnthropic$5.00$25.00$6.82$0.75
34Claude Opus 4.71MAnthropic$5.00$25.00$6.82$0.75
35Claude Opus 4.81MAnthropic$5.00$25.00$6.82$0.75
36Claude Opus 51MAnthropic$5.00$25.00$6.82$0.75
37GPT-5.5256KOpenAI$5.00$30.00$7.27$0.80
38GPT-5.6 Sol258KOpenAI$5.00$30.00$7.27$0.80
39Claude Fable 51MAnthropic$10.00$50.00$13.64$1.50

39 text models · USD per 1M tokens · agent turn = 100K input + 10K output · image models excluded (billed per image-output token)

Cheapest input tokens

What you pay to feed a model context — the number that matters for large-file reads and RAG.

Cheapest output tokens

What you pay for what the model writes — usually the dominant cost on reasoning and code generation.

How this ranking is built

Prices are vendor list rates pulled from the live catalog, the same numbers the API bills against — not a scraped snapshot that goes stale after the next price cut. Every model is scored on one fixed profile so the ordering is comparable, and image-generation models are excluded because they bill per image-output token and would not compare honestly against a chat turn.

Blended $/1M is the same turn expressed per million tokens, which is useful when you want one number to compare against a vendor's headline rate. Both columns move together; neither is a discount RouterPlex applies, because there is no markup to discount.

Price is only half the decision. A cheap model that needs three attempts at a task costs more than an expensive one that lands it first time, so check model benchmarks before moving production traffic to the top of this table.

Questions

What is the cheapest AI model API?

On RouterPlex the cheapest model is DeepSeek V4 Flash at $0.090 per 1M input tokens and $0.18 per 1M output tokens — about $0.011 for a 100K-token-in, 10K-token-out agent turn. Prices are vendor list rates, so the same model costs the same whether you call it here or direct.

Why rank by blended cost instead of input price?

Ranking on input price alone favours models that charge little to read and a lot to write, which is the wrong way round for coding and agent work. This table ranks on a fixed 100K-in / 10K-out turn so input and output rates both count, and the same profile is applied to every model.

How much cheaper is a budget model than a frontier model?

Across this catalog the spread is about 139× on the same agent turn: DeepSeek V4 Flash costs roughly $0.011 where Claude Fable 5 costs about $1.50. Whether that trade is worth it depends on task difficulty — a cheap model that needs three attempts is not cheap.

Is there a cheaper way to buy these models?

Not meaningfully. RouterPlex bills vendor list prices with no markup and no top-up fee, so per-token cost matches going direct to each vendor. What you save is the overhead of separate accounts, keys and minimum balances per provider.

Do cheap models have rate limits or quality caveats?

RouterPlex applies no additional RPM or TPM cap beyond the upstream vendor's. The real caveat is capability: budget models are generally weaker at long agentic tool loops, so compare quality on the benchmarks pages before moving production traffic onto one.

One key reaches every model in this table. Top up from $5 and pay per token.

Cheapest AI Models: Price per 1M Tokens, Ranked · RouterPlex