OpenAI · arena-v1

GPT-5.4 benchmark

Blind preference, cited external scores and RouterPlex evaluation economics. Every result keeps its sample count and source date attached.

Combined quality

91.0

Arena score

Insufficient data

Median latency

Insufficient data

Votes

0

Performance by task.

CategoryQualityArenaExternalEvidence
coding92.7Insufficient data92.7External snapshot
general92.3Insufficient data92.3External snapshot
overall91.0Insufficient data91.0External snapshot
reasoning89.1Insufficient data89.1External snapshot
writing88.0Insufficient data88.0External snapshot

External sources.

LMArena Text Leaderboard · 2026-07-21 text_style_controloverall: rank #35 · 1466.3 raw · 91.0 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

published 2026-07-21 · 62,036 source votes · retrieved 2026-07-22 · LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. LMArena is not affiliated with RouterPlex.

LMArena Text Leaderboard · 2026-07-21 text_style_controlcoding: rank #28 · 1513.8 raw · 92.7 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

published 2026-07-21 · 17,211 source votes · retrieved 2026-07-22 · LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. LMArena is not affiliated with RouterPlex.

LMArena Text Leaderboard · 2026-07-21 text_style_controlreasoning: rank #41 · 1461.4 raw · 89.1 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

published 2026-07-21 · 3,276 source votes · retrieved 2026-07-22 · LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. LMArena is not affiliated with RouterPlex.

LMArena Text Leaderboard · 2026-07-21 text_style_controlwriting: rank #46 · 1433.8 raw · 88.0 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

published 2026-07-21 · 9,993 source votes · retrieved 2026-07-22 · LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. LMArena is not affiliated with RouterPlex.

LMArena Text Leaderboard · 2026-07-21 text_style_controlgeneral: rank #30 · 1461.7 raw · 92.3 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

published 2026-07-21 · 20,856 source votes · retrieved 2026-07-22 · LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. LMArena is not affiliated with RouterPlex.

Compare GPT-5.4.

Test GPT-5.4 yourself.

Use it directly, or put it into a blind head-to-head comparison.

GPT-5.4 Benchmark, Arena Rank and Value · RouterPlex