§ Benchmarks

AI model rankings, with receipts.

Quality signals from blind RouterPlex Arena votes and clearly dated external benchmark snapshots. Price, latency and confidence stay visible instead of disappearing into one unexplained rank.

Current external snapshot

LMArena Text Leaderboard2026-07-21 text_style_control · retrieved 2026-07-22
Rank / modelQualityArenaExternalLatencyEval costVotesConfidence
01Claude Fable 5Anthropic100.0100.00External snapshot
02Claude Opus 4.7Anthropic99.299.20External snapshot
03Claude Opus 4.6Anthropic98.998.90External snapshot
04Kimi K3Moonshot98.198.10External snapshot
05Claude Opus 4.8Anthropic97.697.60External snapshot
06Claude Sonnet 4.6Anthropic97.397.30External snapshot
07Grok 4.5xAI94.994.90External snapshot
08GLM-5.1Zhipu94.494.40External snapshot
09MiMo V2.5 ProXiaomi93.893.80External snapshot
10Kimi K2.6Moonshot93.093.00External snapshot
11GPT-5.4OpenAI92.792.70External snapshot
12Qwen3.7 PlusAlibaba90.690.60External snapshot
13GPT-5.5OpenAI88.788.70External snapshot
14DeepSeek V4 ProDeepSeek86.686.60External snapshot
15MiniMax M3MiniMax85.885.80External snapshot
16Qwen3.6 PlusAlibaba84.184.10External snapshot
17MiMo V2.5Xiaomi82.882.80External snapshot
18DeepSeek V4 FlashDeepSeek80.180.10External snapshot
19MiniMax M2.7MiniMax78.878.80External snapshot

How ranking works.

External results retain their raw rating, source rank, confidence interval, vote count, version and retrieval date. RouterPlex converts each source rank to a within-category percentile so different category populations stay on a consistent 0–100 display scale.

How Arena joins later.

RouterPlex Arena uses versioned blind pairwise votes. Its weight grows only with valid sample confidence, while custom prompts remain private and never affect the public ranking.

Run a blind comparison