
Benchmarks
Rank the signal, inspect the evidence.
Blind RouterPlex preferences and dated external snapshots, shown beside latency, evaluation cost, sample size, and confidence.
20
Ranked models
1
Cited sources
Ranking console
Published leaderboard.
Select a task and scoring signal. Every row keeps the underlying evidence visible.
| Rank / model | External signal | Arena | External | Latency | Eval cost | Votes | Evidence |
|---|---|---|---|---|---|---|---|
01 Claude Fable 5 Anthropic | 100.0 | - | 100.0 | - | - | 0 | Cited |
02 Qwen3.8 Max Alibaba | 99.5 | - | 99.5 | - | - | 0 | Cited |
03 Claude Opus 4.7 Anthropic | 98.7 | - | 98.7 | - | - | 0 | Cited |
04 Claude Opus 4.6 Anthropic | 98.2 | - | 98.2 | - | - | 0 | Cited |
05 Claude Opus 4.8 Anthropic | 95.5 | 51.2 | 95.5 | 11889ms | $0.0063 | 1 | low |
06 Claude Sonnet 4.6 Anthropic | 93.7 | - | 93.7 | - | - | 0 | Cited |
07 GLM-5.1 Zhipu | 93.2 | - | 93.2 | - | - | 0 | Cited |
08 Grok 4.5 xAI | 92.6 | - | 92.6 | - | - | 0 | Cited |
09 GPT-5.5 OpenAI | 92.1 | - | 92.1 | - | - | 0 | Cited |
10 DeepSeek V4 Pro DeepSeek | 90.5 | - | 90.5 | - | - | 0 | Cited |
11 Qwen3.7 Plus Alibaba | 88.7 | - | 88.7 | - | - | 0 | Cited |
12 GPT-5.4 OpenAI | 87.6 | - | 87.6 | - | - | 0 | Cited |
13 MiMo V2.5 Pro Xiaomi | 86.8 | - | 86.8 | - | - | 0 | Cited |
14 Kimi K2.6 Moonshot | 86.3 | - | 86.3 | - | - | 0 | Cited |
15 Hy3 Tencent Hunyuan | 84.5 | - | 84.5 | - | - | 0 | Cited |
16 DeepSeek V4 Flash DeepSeek | 81.1 | - | 81.1 | - | - | 0 | Cited |
17 Qwen3.6 Plus Alibaba | 80.5 | 48.8 | 80.5 | 20978ms | $0.0033 | 1 | low |
18 MiniMax M3 MiniMax | 79.5 | - | 79.5 | - | - | 0 | Cited |
19 MiMo V2.5 Xiaomi | 74.7 | - | 74.7 | - | - | 0 | Cited |
20 MiniMax M2.7 MiniMax | 63.7 | - | 63.7 | - | - | 0 | Cited |
External snapshot
Source and retrieval date stay attached to every imported signal.
Method / ranking-v1
Transparent by construction.
Normalize, never obscure
External results retain their raw rating, source rank, interval, vote count, version, and retrieval date. Source rank becomes a within-category percentile on the common 0-100 scale.
Confidence earns its weight
Blind Arena votes grow in influence only as valid sample confidence improves. Private prompts stay private and never enter the public ranking.