
Benchmarks
Rank the signal, inspect the evidence.
Blind RouterPlex preferences and dated external snapshots, shown beside latency, evaluation cost, sample size, and confidence.
20
Ranked models
1
Cited sources
Ranking console
Published leaderboard.
Select a task and scoring signal. Every row keeps the underlying evidence visible.
| Rank / model | Quality signal | Arena | External | Latency | Eval cost | Votes | Evidence |
|---|---|---|---|---|---|---|---|
01 Claude Fable 5 Anthropic | 100.0 | - | 100.0 | - | - | 0 | Cited |
02 Claude Opus 4.6 Anthropic | 99.2 | - | 99.2 | - | - | 0 | Cited |
03 Qwen3.8 Max Alibaba | 99.0 | - | 99.0 | - | - | 0 | Cited |
04 Claude Opus 4.7 Anthropic | 98.7 | - | 98.7 | - | - | 0 | Cited |
05 GPT-5.5 OpenAI | 95.5 | - | 95.5 | - | - | 0 | Cited |
06 Claude Sonnet 4.6 Anthropic | 92.7 | - | 92.7 | - | - | 0 | Cited |
07 Grok 4.5 xAI | 91.6 | - | 91.6 | - | - | 0 | Cited |
08 Claude Opus 4.8 Anthropic | 91.5 | 51.2 | 94.5 | 11889ms | $0.0063 | 1 | low |
09 GLM-5.1 Zhipu | 91.4 | - | 91.4 | - | - | 0 | Cited |
10 MiMo V2.5 Pro Xiaomi | 90.3 | - | 90.3 | - | - | 0 | Cited |
11 GPT-5.4 OpenAI | 90.1 | - | 90.1 | - | - | 0 | Cited |
12 Kimi K2.6 Moonshot | 89.3 | - | 89.3 | - | - | 0 | Cited |
13 Qwen3.7 Plus Alibaba | 88.5 | - | 88.5 | - | - | 0 | Cited |
14 DeepSeek V4 Pro DeepSeek | 88.0 | - | 88.0 | - | - | 0 | Cited |
15 Hy3 Tencent Hunyuan | 87.2 | - | 87.2 | - | - | 0 | Cited |
16 MiniMax M3 MiniMax | 82.7 | - | 82.7 | - | - | 0 | Cited |
17 Qwen3.6 Plus Alibaba | 79.9 | 48.8 | 82.2 | 20978ms | $0.0033 | 1 | low |
18 DeepSeek V4 Flash DeepSeek | 79.6 | - | 79.6 | - | - | 0 | Cited |
19 MiMo V2.5 Xiaomi | 78.3 | - | 78.3 | - | - | 0 | Cited |
20 MiniMax M2.7 MiniMax | 70.9 | - | 70.9 | - | - | 0 | Cited |
External snapshot
Source and retrieval date stay attached to every imported signal.
Method / ranking-v1
Transparent by construction.
Normalize, never obscure
External results retain their raw rating, source rank, interval, vote count, version, and retrieval date. Source rank becomes a within-category percentile on the common 0-100 scale.
Confidence earns its weight
Blind Arena votes grow in influence only as valid sample confidence improves. Private prompts stay private and never enter the public ranking.