
Benchmarks
Rank the signal, inspect the evidence.
Blind RouterPlex preferences and dated external snapshots, shown beside latency, evaluation cost, sample size, and confidence.
18
Ranked models
1
Cited sources
Ranking console
Published leaderboard.
Select a task and scoring signal. Every row keeps the underlying evidence visible.
| Rank / model | External signal | Arena | External | Latency | Eval cost | Votes | Evidence |
|---|---|---|---|---|---|---|---|
01 Claude Fable 5 Anthropic | 99.5 | - | 99.5 | - | - | 0 | Cited |
02 Claude Opus 4.6 Anthropic | 98.6 | - | 98.6 | - | - | 0 | Cited |
03 GPT-5.5 OpenAI | 97.8 | - | 97.8 | - | - | 0 | Cited |
04 Claude Opus 4.7 Anthropic | 96.7 | - | 96.7 | - | - | 0 | Cited |
05 Grok 4.5 xAI | 95.7 | - | 95.7 | - | - | 0 | Cited |
06 GLM-5.1 Zhipu | 95.1 | - | 95.1 | - | - | 0 | Cited |
07 Kimi K2.6 Moonshot | 94.9 | - | 94.9 | - | - | 0 | Cited |
08 MiMo V2.5 Pro Xiaomi | 92.7 | - | 92.7 | - | - | 0 | Cited |
09 Claude Opus 4.8 Anthropic | 91.9 | - | 91.9 | - | - | 0 | Cited |
10 Qwen3.7 Plus Alibaba | 90.8 | - | 90.8 | - | - | 0 | Cited |
11 Claude Sonnet 4.6 Anthropic | 88.9 | - | 88.9 | - | - | 0 | Cited |
12 GPT-5.4 OpenAI | 88.1 | - | 88.1 | - | - | 0 | Cited |
13 Qwen3.6 Plus Alibaba | 87.3 | - | 87.3 | - | - | 0 | Cited |
14 DeepSeek V4 Pro DeepSeek | 84.0 | - | 84.0 | - | - | 0 | Cited |
15 MiMo V2.5 Xiaomi | 82.7 | - | 82.7 | - | - | 0 | Cited |
16 MiniMax M3 MiniMax | 81.6 | - | 81.6 | - | - | 0 | Cited |
17 DeepSeek V4 Flash DeepSeek | 76.2 | - | 76.2 | - | - | 0 | Cited |
18 MiniMax M2.7 MiniMax | 74.3 | - | 74.3 | - | - | 0 | Cited |
External snapshot
Source and retrieval date stay attached to every imported signal.
Method / ranking-v1
Transparent by construction.
Normalize, never obscure
External results retain their raw rating, source rank, interval, vote count, version, and retrieval date. Source rank becomes a within-category percentile on the common 0-100 scale.
Confidence earns its weight
Blind Arena votes grow in influence only as valid sample confidence improves. Private prompts stay private and never enter the public ranking.