§ Benchmarks

AI model rankings, with receipts.

Quality signals from blind RouterPlex Arena votes and clearly dated external benchmark snapshots. Price, latency and confidence stay visible instead of disappearing into one unexplained rank.

Current external snapshot

LMArena Text Leaderboard2026-07-21 text_style_control · retrieved 2026-07-22
Rank / modelQualityArenaExternalLatencyEval costVotesConfidence
01Claude Fable 5Anthropic100.0100.00External snapshot
02Claude Opus 4.6Anthropic99.299.20External snapshot
03Claude Opus 4.7Anthropic98.798.70External snapshot
04Kimi K3Moonshot97.197.10External snapshot
05Claude Sonnet 4.6Anthropic96.896.80External snapshot
06Claude Opus 4.8Anthropic96.396.30External snapshot
07GPT-5.5OpenAI95.295.20External snapshot
08MiMo V2.5 ProXiaomi95.095.00External snapshot
09GLM-5.1Zhipu93.693.60External snapshot
10GPT-5.4OpenAI92.392.30External snapshot
11Grok 4.5xAI92.092.00External snapshot
12Kimi K2.6Moonshot90.290.20External snapshot
13DeepSeek V4 ProDeepSeek89.989.90External snapshot
14Qwen3.7 PlusAlibaba89.189.10External snapshot
15MiniMax M3MiniMax85.185.10External snapshot
16Qwen3.6 PlusAlibaba84.684.60External snapshot
17MiMo V2.5Xiaomi82.882.80External snapshot
18DeepSeek V4 FlashDeepSeek81.781.70External snapshot
19MiniMax M2.7MiniMax73.573.50External snapshot

How ranking works.

External results retain their raw rating, source rank, confidence interval, vote count, version and retrieval date. RouterPlex converts each source rank to a within-category percentile so different category populations stay on a consistent 0–100 display scale.

How Arena joins later.

RouterPlex Arena uses versioned blind pairwise votes. Its weight grows only with valid sample confidence, while custom prompts remain private and never affect the public ranking.

Run a blind comparison