OpenAI · arena-v1

GPT-5.5 benchmark

Blind preference, cited external scores and RouterPlex evaluation economics. Every result keeps its sample count and source date attached.

Combined quality

96.0

Arena score

Insufficient data

Median latency

Insufficient data

Votes

0

Performance by task.

CategoryQualityArenaExternalEvidence
coding88.7Insufficient data88.7External snapshot
general95.2Insufficient data95.2External snapshot
overall96.0Insufficient data96.0External snapshot
reasoning98.1Insufficient data98.1External snapshot
writing92.5Insufficient data92.5External snapshot

External sources.

LMArena Text Leaderboard · 2026-07-21 text_style_controloverall: rank #16 · 1476.3 raw · 96.0 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

published 2026-07-21 · 47,180 source votes · retrieved 2026-07-22 · LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. LMArena is not affiliated with RouterPlex.

LMArena Text Leaderboard · 2026-07-21 text_style_controlcoding: rank #43 · 1507.2 raw · 88.7 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

published 2026-07-21 · 13,423 source votes · retrieved 2026-07-22 · LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. LMArena is not affiliated with RouterPlex.

LMArena Text Leaderboard · 2026-07-21 text_style_controlreasoning: rank #8 · 1497.1 raw · 98.1 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

published 2026-07-21 · 2,436 source votes · retrieved 2026-07-22 · LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. LMArena is not affiliated with RouterPlex.

LMArena Text Leaderboard · 2026-07-21 text_style_controlwriting: rank #29 · 1447.0 raw · 92.5 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

published 2026-07-21 · 8,118 source votes · retrieved 2026-07-22 · LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. LMArena is not affiliated with RouterPlex.

LMArena Text Leaderboard · 2026-07-21 text_style_controlgeneral: rank #19 · 1471.5 raw · 95.2 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

published 2026-07-21 · 15,888 source votes · retrieved 2026-07-22 · LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. LMArena is not affiliated with RouterPlex.

Compare GPT-5.5.

Test GPT-5.5 yourself.

Use it directly, or put it into a blind head-to-head comparison.

GPT-5.5 Benchmark, Arena Rank and Value · RouterPlex