
Benchmarks
Rank the signal, inspect the evidence.
Blind RouterPlex preferences and dated external snapshots, shown beside latency, evaluation cost, sample size, and confidence.
20
Ranked models
1
Cited sources
Ranking console
Published leaderboard.
Select a task and scoring signal. Every row keeps the underlying evidence visible.
| Rank / model | Quality signal | Arena | External | Latency | Eval cost | Votes | Evidence |
|---|---|---|---|---|---|---|---|
01 Claude Fable 5 Anthropic | 100.0 | - | 100.0 | - | - | 0 | Cited |
02 Claude Opus 4.6 Anthropic | 99.2 | - | 99.2 | - | - | 0 | Cited |
03 Claude Opus 4.7 Anthropic | 98.4 | - | 98.4 | - | - | 0 | Cited |
04 Qwen3.8 Max Alibaba | 97.9 | - | 97.9 | - | - | 0 | Cited |
05 Claude Opus 4.8 Anthropic | 96.6 | - | 96.6 | - | - | 0 | Cited |
06 Claude Sonnet 4.6 Anthropic | 96.3 | - | 96.3 | - | - | 0 | Cited |
07 GPT-5.5 OpenAI | 94.8 | - | 94.8 | - | - | 0 | Cited |
08 MiMo V2.5 Pro Xiaomi | 94.0 | - | 94.0 | - | - | 0 | Cited |
09 GLM-5.1 Zhipu | 93.5 | - | 93.5 | - | - | 0 | Cited |
10 Grok 4.5 xAI | 91.9 | - | 91.9 | - | - | 0 | Cited |
11 GPT-5.4 OpenAI | 91.1 | - | 91.1 | - | - | 0 | Cited |
12 DeepSeek V4 Pro DeepSeek | 89.5 | - | 89.5 | - | - | 0 | Cited |
13 Kimi K2.6 Moonshot | 89.3 | - | 89.3 | - | - | 0 | Cited |
14 Qwen3.7 Plus Alibaba | 88.5 | - | 88.5 | - | - | 0 | Cited |
15 Hy3 Tencent Hunyuan | 86.9 | - | 86.9 | - | - | 0 | Cited |
16 MiniMax M3 MiniMax | 84.8 | - | 84.8 | - | - | 0 | Cited |
17 Qwen3.6 Plus Alibaba | 83.5 | - | 83.5 | - | - | 0 | Cited |
18 MiMo V2.5 Xiaomi | 81.9 | - | 81.9 | - | - | 0 | Cited |
19 DeepSeek V4 Flash DeepSeek | 80.9 | - | 80.9 | - | - | 0 | Cited |
20 MiniMax M2.7 MiniMax | 72.5 | - | 72.5 | - | - | 0 | Cited |
External snapshot
Source and retrieval date stay attached to every imported signal.
Method / ranking-v1
Transparent by construction.
Normalize, never obscure
External results retain their raw rating, source rank, interval, vote count, version, and retrieval date. Source rank becomes a within-category percentile on the common 0-100 scale.
Confidence earns its weight
Blind Arena votes grow in influence only as valid sample confidence improves. Private prompts stay private and never enter the public ranking.