§ Benchmarks

AI model rankings, with receipts.

Quality signals from blind RouterPlex Arena votes and clearly dated external benchmark snapshots. Price, latency and confidence stay visible instead of disappearing into one unexplained rank.

Current external snapshot

LMArena Text Leaderboard2026-07-21 text_style_control · retrieved 2026-07-22
Rank / modelQualityArenaExternalLatencyEval costVotesConfidence
01Claude Fable 5Anthropic100.0100.00External snapshot
02Claude Opus 4.6Anthropic98.698.60External snapshot
03Grok 4.5xAI98.498.40External snapshot
04GPT-5.5OpenAI98.198.10External snapshot
05Claude Opus 4.7Anthropic97.397.30External snapshot
06Kimi K2.6Moonshot95.495.40External snapshot
07GLM-5.1Zhipu95.195.10External snapshot
08Claude Opus 4.8Anthropic94.394.30External snapshot
09MiMo V2.5 ProXiaomi93.293.20External snapshot
10Qwen3.7 PlusAlibaba92.692.60External snapshot
11Claude Sonnet 4.6Anthropic89.989.90External snapshot
12GPT-5.4OpenAI89.189.10External snapshot
13Qwen3.6 PlusAlibaba87.787.70External snapshot
14DeepSeek V4 ProDeepSeek85.085.00External snapshot
15MiniMax M3MiniMax83.983.90External snapshot
16MiMo V2.5Xiaomi83.183.10External snapshot
17DeepSeek V4 FlashDeepSeek76.076.00External snapshot
18MiniMax M2.7MiniMax75.175.10External snapshot

How ranking works.

External results retain their raw rating, source rank, confidence interval, vote count, version and retrieval date. RouterPlex converts each source rank to a within-category percentile so different category populations stay on a consistent 0–100 display scale.

How Arena joins later.

RouterPlex Arena uses versioned blind pairwise votes. Its weight grows only with valid sample confidence, while custom prompts remain private and never affect the public ranking.

Run a blind comparison