§ Benchmarks

AI model rankings, with receipts.

Quality signals from blind RouterPlex Arena votes and clearly dated external benchmark snapshots. Price, latency and confidence stay visible instead of disappearing into one unexplained rank.

Current external snapshot

LMArena Text Leaderboard2026-07-21 text_style_control · retrieved 2026-07-22
Rank / modelQualityArenaExternalLatencyEval costVotesConfidence
01Claude Fable 5Anthropic100.0100.00External snapshot
02Claude Opus 4.7Anthropic98.998.90External snapshot
03Claude Opus 4.6Anthropic98.498.40External snapshot
04Kimi K3Moonshot97.697.60External snapshot
05Claude Opus 4.8Anthropic95.595.50External snapshot
06GLM-5.1Zhipu94.194.10External snapshot
07Claude Sonnet 4.6Anthropic93.993.90External snapshot
08GPT-5.5OpenAI92.592.50External snapshot
09DeepSeek V4 ProDeepSeek89.989.90External snapshot
10Grok 4.5xAI89.389.30External snapshot
11Qwen3.7 PlusAlibaba89.189.10External snapshot
12GPT-5.4OpenAI88.088.00External snapshot
13MiMo V2.5 ProXiaomi87.787.70External snapshot
14Kimi K2.6Moonshot86.486.40External snapshot
15Qwen3.6 PlusAlibaba81.981.90External snapshot
16DeepSeek V4 FlashDeepSeek81.681.60External snapshot
17MiniMax M3MiniMax81.381.30External snapshot
18MiMo V2.5Xiaomi74.474.40External snapshot
19MiniMax M2.7MiniMax64.564.50External snapshot

How ranking works.

External results retain their raw rating, source rank, confidence interval, vote count, version and retrieval date. RouterPlex converts each source rank to a within-category percentile so different category populations stay on a consistent 0–100 display scale.

How Arena joins later.

RouterPlex Arena uses versioned blind pairwise votes. Its weight grows only with valid sample confidence, while custom prompts remain private and never affect the public ranking.

Run a blind comparison