
Benchmarks
Rank the signal, inspect the evidence.
Blind RouterPlex preferences and dated external snapshots, shown beside latency, evaluation cost, sample size, and confidence.
20
Ranked models
1
Cited sources
Ranking console
Published leaderboard.
Select a task and scoring signal. Every row keeps the underlying evidence visible.
| Rank / model | Quality signal | Arena | External | Latency | Eval cost | Votes | Evidence |
|---|---|---|---|---|---|---|---|
01 Claude Fable 5 Anthropic | 100.0 | - | 100.0 | - | - | 0 | Cited |
02 Claude Opus 4.6 Anthropic | 99.2 | - | 99.2 | - | - | 0 | Cited |
03 Claude Opus 4.7 Anthropic | 98.9 | - | 98.9 | - | - | 0 | Cited |
04 Qwen3.8 Max Alibaba | 97.9 | - | 97.9 | - | - | 0 | Cited |
05 Claude Opus 4.8 Anthropic | 97.3 | - | 97.3 | - | - | 0 | Cited |
06 Claude Sonnet 4.6 Anthropic | 96.8 | - | 96.8 | - | - | 0 | Cited |
07 Grok 4.5 xAI | 93.9 | - | 93.9 | - | - | 0 | Cited |
08 GLM-5.1 Zhipu | 93.1 | - | 93.1 | - | - | 0 | Cited |
09 MiMo V2.5 Pro Xiaomi | 92.8 | - | 92.8 | - | - | 0 | Cited |
10 Kimi K2.6 Moonshot | 92.3 | - | 92.3 | - | - | 0 | Cited |
11 GPT-5.4 OpenAI | 91.8 | - | 91.8 | - | - | 0 | Cited |
12 Qwen3.7 Plus Alibaba | 89.7 | - | 89.7 | - | - | 0 | Cited |
13 GPT-5.5 OpenAI | 89.1 | - | 89.1 | - | - | 0 | Cited |
14 DeepSeek V4 Pro DeepSeek | 86.5 | - | 86.5 | - | - | 0 | Cited |
15 Hy3 Tencent Hunyuan | 85.4 | - | 85.4 | - | - | 0 | Cited |
16 MiniMax M3 MiniMax | 84.4 | - | 84.4 | - | - | 0 | Cited |
17 Qwen3.6 Plus Alibaba | 83.0 | - | 83.0 | - | - | 0 | Cited |
18 MiMo V2.5 Xiaomi | 82.2 | - | 82.2 | - | - | 0 | Cited |
19 DeepSeek V4 Flash DeepSeek | 79.0 | - | 79.0 | - | - | 0 | Cited |
20 MiniMax M2.7 MiniMax | 77.7 | - | 77.7 | - | - | 0 | Cited |
External snapshot
Source and retrieval date stay attached to every imported signal.
Method / ranking-v1
Transparent by construction.
Normalize, never obscure
External results retain their raw rating, source rank, interval, vote count, version, and retrieval date. Source rank becomes a within-category percentile on the common 0-100 scale.
Confidence earns its weight
Blind Arena votes grow in influence only as valid sample confidence improves. Private prompts stay private and never enter the public ranking.