Benchmarks

Rank the signal, inspect the evidence.

Blind RouterPlex preferences and dated external snapshots, shown beside latency, evaluation cost, sample size, and confidence.

Ranking console

Published leaderboard.

Select a task and scoring signal. Every row keeps the underlying evidence visible.

Rank / modelQuality signalArenaExternalLatencyEval costVotesEvidence
01
100.0
-100.0--0Cited
02
99.2
-99.2--0Cited
03
98.4
-98.4--0Cited
04
97.9
-97.9--0Cited
05
96.6
-96.6--0Cited
96.3
-96.3--0Cited
07
GPT-5.5

OpenAI

94.8
-94.8--0Cited
94.0
-94.0--0Cited
09
GLM-5.1

Zhipu

93.5
-93.5--0Cited
91.9
-91.9--0Cited
11
GPT-5.4

OpenAI

91.1
-91.1--0Cited
12
89.5
-89.5--0Cited
13
Kimi K2.6

Moonshot

89.3
-89.3--0Cited
14
88.5
-88.5--0Cited
15
Hy3

Tencent Hunyuan

86.9
-86.9--0Cited
16
MiniMax M3

MiniMax

84.8
-84.8--0Cited
17
83.5
-83.5--0Cited
18
MiMo V2.5

Xiaomi

81.9
-81.9--0Cited
80.9
-80.9--0Cited
20
72.5
-72.5--0Cited
0-100 normalized display scalegeneral / combined quality

External snapshot

Source and retrieval date stay attached to every imported signal.

LMArena Text Leaderboard2026-08-03 text_style_control 0fdff69b4fd7 / retrieved 2026-08-07

Method / ranking-v1

Transparent by construction.

Normalize, never obscure

External results retain their raw rating, source rank, interval, vote count, version, and retrieval date. Source rank becomes a within-category percentile on the common 0-100 scale.

Confidence earns its weight

Blind Arena votes grow in influence only as valid sample confidence improves. Private prompts stay private and never enter the public ranking.

Run a blind comparison