Benchmarks

Rank the signal, inspect the evidence.

Blind RouterPlex preferences and dated external snapshots, shown beside latency, evaluation cost, sample size, and confidence.

Ranking console

Published leaderboard.

Select a task and scoring signal. Every row keeps the underlying evidence visible.

Rank / modelQuality signalArenaExternalLatencyEval costVotesEvidence
01
100.0
-100.0--0Cited
02
99.2
-99.2--0Cited
03
98.9
-98.9--0Cited
04
97.9
-97.9--0Cited
05
97.3
-97.3--0Cited
96.8
-96.8--0Cited
93.9
-93.9--0Cited
08
GLM-5.1

Zhipu

93.1
-93.1--0Cited
92.8
-92.8--0Cited
10
Kimi K2.6

Moonshot

92.3
-92.3--0Cited
11
GPT-5.4

OpenAI

91.8
-91.8--0Cited
12
89.7
-89.7--0Cited
13
GPT-5.5

OpenAI

89.1
-89.1--0Cited
14
86.5
-86.5--0Cited
15
Hy3

Tencent Hunyuan

85.4
-85.4--0Cited
16
MiniMax M3

MiniMax

84.4
-84.4--0Cited
17
83.0
-83.0--0Cited
18
MiMo V2.5

Xiaomi

82.2
-82.2--0Cited
79.0
-79.0--0Cited
20
77.7
-77.7--0Cited
0-100 normalized display scalecoding / combined quality

External snapshot

Source and retrieval date stay attached to every imported signal.

LMArena Text Leaderboard2026-08-03 text_style_control 0fdff69b4fd7 / retrieved 2026-08-07

Method / ranking-v1

Transparent by construction.

Normalize, never obscure

External results retain their raw rating, source rank, interval, vote count, version, and retrieval date. Source rank becomes a within-category percentile on the common 0-100 scale.

Confidence earns its weight

Blind Arena votes grow in influence only as valid sample confidence improves. Private prompts stay private and never enter the public ranking.

Run a blind comparison