§ Benchmarks

AI model rankings, with receipts.

Quality signals from blind RouterPlex Arena votes and clearly dated external benchmark snapshots. Price, latency and confidence stay visible instead of disappearing into one unexplained rank.

Current external snapshot

LMArena Text Leaderboard2026-07-21 text_style_control · retrieved 2026-07-22
Rank / modelQualityArenaExternalLatencyEval costVotesConfidence
01Claude Fable 5Anthropic100.0100.00External snapshot
02Claude Opus 4.6Anthropic99.299.20External snapshot
03Claude Opus 4.7Anthropic98.798.70External snapshot
04Kimi K3Moonshot97.697.60External snapshot
05GPT-5.5OpenAI96.096.00External snapshot
06Claude Opus 4.8Anthropic94.494.40External snapshot
07Claude Sonnet 4.6Anthropic93.193.10External snapshot
08GLM-5.1Zhipu92.692.60External snapshot
09Grok 4.5xAI91.591.50External snapshot
10MiMo V2.5 ProXiaomi91.291.20External snapshot
11GPT-5.4OpenAI91.091.00External snapshot
12Kimi K2.6Moonshot89.989.90External snapshot
13Qwen3.7 PlusAlibaba89.489.40External snapshot
14DeepSeek V4 ProDeepSeek88.188.10External snapshot
15MiniMax M3MiniMax83.383.30External snapshot
16Qwen3.6 PlusAlibaba83.083.00External snapshot
17DeepSeek V4 FlashDeepSeek80.180.10External snapshot
18MiMo V2.5Xiaomi79.079.00External snapshot
19MiniMax M2.7MiniMax71.671.60External snapshot

How ranking works.

External results retain their raw rating, source rank, confidence interval, vote count, version and retrieval date. RouterPlex converts each source rank to a within-category percentile so different category populations stay on a consistent 0–100 display scale.

How Arena joins later.

RouterPlex Arena uses versioned blind pairwise votes. Its weight grows only with valid sample confidence, while custom prompts remain private and never affect the public ranking.

Run a blind comparison