Benchmarks

Rank the signal, inspect the evidence.

Blind RouterPlex preferences and dated external snapshots, shown beside latency, evaluation cost, sample size, and confidence.

Ranking console

Published leaderboard.

Select a task and scoring signal. Every row keeps the underlying evidence visible.

Rank / modelQuality signalArenaExternalLatencyEval costVotesEvidence
01
99.5
-99.5--0Cited
02
98.6
-98.6--0Cited
03
GPT-5.5

OpenAI

97.8
-97.8--0Cited
04
96.7
-96.7--0Cited
95.7
-95.7--0Cited
06
GLM-5.1

Zhipu

95.1
-95.1--0Cited
07
Kimi K2.6

Moonshot

94.9
-94.9--0Cited
92.7
-92.7--0Cited
09
91.9
-91.9--0Cited
10
90.8
-90.8--0Cited
88.9
-88.9--0Cited
12
GPT-5.4

OpenAI

88.1
-88.1--0Cited
13
87.3
-87.3--0Cited
14
84.0
-84.0--0Cited
15
MiMo V2.5

Xiaomi

82.7
-82.7--0Cited
16
MiniMax M3

MiniMax

81.6
-81.6--0Cited
76.2
-76.2--0Cited
18
74.3
-74.3--0Cited
0-100 normalized display scalereasoning / combined quality

External snapshot

Source and retrieval date stay attached to every imported signal.

LMArena Text Leaderboard2026-08-03 text_style_control 0fdff69b4fd7 / retrieved 2026-08-07

Method / ranking-v1

Transparent by construction.

Normalize, never obscure

External results retain their raw rating, source rank, interval, vote count, version, and retrieval date. Source rank becomes a within-category percentile on the common 0-100 scale.

Confidence earns its weight

Blind Arena votes grow in influence only as valid sample confidence improves. Private prompts stay private and never enter the public ranking.

Run a blind comparison