All benchmarks

OpenAI / arena-v1

GPT-5.5 benchmark.

Blind preference, cited external scores, and evaluation economics. Every result keeps its sample count and source date attached.

Evidence levelCited snapshot
Last update2026-08-07
Categories5

Combined quality

95.5

Arena score

Pending

Median latency

Pending

Valid votes

0

Task analysis

Performance by task.

CategoryQualityArenaExternalEvidence
coding89.1Pending89.1Cited snapshot
general94.8Pending94.8Cited snapshot
overall95.5Pending95.5Cited snapshot
reasoning97.8Pending97.8Cited snapshot
writing92.1Pending92.1Cited snapshot

Provenance

External evidence.

LMArena Text Leaderboard

2026-08-03 text_style_control 0fdff69b4fd7

overallRank #181476.3 raw95.5 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

Published 2026-08-03 / 51,172 source votes / retrieved 2026-08-07 / LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. Official leaderboard: https://arena.ai/leaderboard. Dataset: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset. Dataset commit: 4e52c8e709c90a4cad8498d9db5aad11709b04e0. LMArena is not affiliated with RouterPlex.

LMArena Text Leaderboard

2026-08-03 text_style_control 0fdff69b4fd7

codingRank #421508.5 raw89.1 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

Published 2026-08-03 / 14,576 source votes / retrieved 2026-08-07 / LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. Official leaderboard: https://arena.ai/leaderboard. Dataset: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset. Dataset commit: 4e52c8e709c90a4cad8498d9db5aad11709b04e0. LMArena is not affiliated with RouterPlex.

LMArena Text Leaderboard

2026-08-03 text_style_control 0fdff69b4fd7

reasoningRank #91499.3 raw97.8 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

Published 2026-08-03 / 2,592 source votes / retrieved 2026-08-07 / LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. Official leaderboard: https://arena.ai/leaderboard. Dataset: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset. Dataset commit: 4e52c8e709c90a4cad8498d9db5aad11709b04e0. LMArena is not affiliated with RouterPlex.

LMArena Text Leaderboard

2026-08-03 text_style_control 0fdff69b4fd7

writingRank #311447.3 raw92.1 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

Published 2026-08-03 / 8,923 source votes / retrieved 2026-08-07 / LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. Official leaderboard: https://arena.ai/leaderboard. Dataset: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset. Dataset commit: 4e52c8e709c90a4cad8498d9db5aad11709b04e0. LMArena is not affiliated with RouterPlex.

LMArena Text Leaderboard

2026-08-03 text_style_control 0fdff69b4fd7

generalRank #211472.7 raw94.8 percentile

Bradley-Terry ratings from LMArena's style-controlled text leaderboard. RouterPlex preserves the raw rating, source rank, confidence interval, vote count, and publish date. The displayed external score is the source-rank percentile within each source category.

Published 2026-08-03 / 17,288 source votes / retrieved 2026-08-07 / LMArena leaderboard dataset, licensed CC BY 4.0. RouterPlex includes only exact case-insensitive model-ID matches. Official leaderboard: https://arena.ai/leaderboard. Dataset: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset. Dataset commit: 4e52c8e709c90a4cad8498d9db5aad11709b04e0. LMArena is not affiliated with RouterPlex.

Compare GPT-5.5

Your workload is the final benchmark

Test GPT-5.5 yourself.

Use it directly, or put it into a blind head-to-head comparison.

GPT-5.5 Benchmarks · RouterPlex