Benchmarks / Human Preference
LMArena Text
Head-to-head human preference votes, Bradley–Terry Arena Score (Elo-like, style-controlled by default). Pool-relative: never comparable across arenas.
maintained by LMArena (Arena) · arena.ai · 135,940 observations on record
Versions
29 versions
| Version | Tasks |
|---|---|
| Overall (style control) current lmarena-text-overall | — |
| Chinese (style control) lmarena-text-chinese | — |
| Coding (style control) lmarena-text-coding | — |
| Creative Writing (style control) lmarena-text-creative-writing | — |
| English (style control) lmarena-text-english | — |
| Exclude Ties (style control) lmarena-text-exclude-ties | — |
| Expert (style control) lmarena-text-expert | — |
| French (style control) lmarena-text-french | — |
| German (style control) lmarena-text-german | — |
| Hard Prompts (style control) lmarena-text-hard-prompts | — |
| Hard Prompts English (style control) lmarena-text-hard-prompts-english | — |
| Industry Business And Management And Financial Operations (style control) lmarena-text-industry-business-and-management-and-financial-operations | — |
| Industry Entertainment And Sports And Media (style control) lmarena-text-industry-entertainment-and-sports-and-media | — |
| Industry Legal And Government (style control) lmarena-text-industry-legal-and-government | — |
| Industry Life And Physical And Social Science (style control) lmarena-text-industry-life-and-physical-and-social-science | — |
| Industry Mathematical (style control) lmarena-text-industry-mathematical | — |
| Industry Medicine And Healthcare (style control) lmarena-text-industry-medicine-and-healthcare | — |
| Industry Software And It Services (style control) lmarena-text-industry-software-and-it-services | — |
| Industry Writing And Literature And Language (style control) lmarena-text-industry-writing-and-literature-and-language | — |
| Instruction Following (style control) lmarena-text-instruction-following | — |
| Japanese (style control) lmarena-text-japanese | — |
| Korean (style control) lmarena-text-korean | — |
| Longer Query (style control) lmarena-text-longer-query | — |
| Math (style control) lmarena-text-math | — |
| Multi Turn (style control) lmarena-text-multi-turn | — |
| Non English (style control) lmarena-text-non-english | — |
| Polish (style control) lmarena-text-polish | — |
| Russian (style control) lmarena-text-russian | — |
| Spanish (style control) lmarena-text-spanish | — |
Current leaderboard
Overall (style control) · Arena score (Elo) · top 25 of 384 systems
| # | System | Arena score (Elo) |
|---|---|---|
| 1 | Claude Fable 5 Anthropic | 1507.2 1502.2–1512.1 |
| 2 | Claude Opus 4.6 Thinking Anthropic | 1504.7 1501.2–1508.2 |
| 3 | Claude Opus 4.7 Thinking Anthropic | 1502.3 1498.4–1506.2 |
| 4 | Claude Opus 4.6 Anthropic | 1497.6 1494.1–1501.0 |
| 5 | Claude Opus 4.7 Anthropic | 1494.5 1490.6–1498.4 |
| 6 | Muse Spark 1.1 Meta AI | 1492.1 1487.0–1497.1 |
| 7 | Muse Spark Meta AI | 1488.2 1482.3–1494.0 |
| 8 | Gemini 3.1 Pro Preview Google DeepMind | 1486.7 1483.5–1489.9 |
| 9 | Gemini 3 Pro Google | 1485.6 1481.7–1489.4 |
| 10 | GPT 5.6 Solxhigh effort OpenAI | 1484.6 1473.2–1496.1 |
| 11 | GPT 5.6 Sol OpenAI | 1483.2 1478.0–1488.3 |
| 12 | GPT 5.5 (2026-04-22)high effort OpenAI | 1481.9 1477.3–1486.4 |
| 13 | Claude Opus 4.8 Thinking Anthropic | 1481.2 1476.6–1485.7 |
| 14 | GPT 5.5 (2026-04-22) OpenAI | 1477.2 1473.4–1481.0 |
| 15 | GPT 5.4 (2026-03-05)high effort OpenAI | 1476.6 1472.5–1480.6 |
| 16 | Gemini 3.5 Flashhigh effort Google DeepMind | 1476.5 1469.9–1483.0 |
| 17 | GPT 5.2 Chat Latest 20260210 OpenAI | 1476.3 1472.2–1480.4 |
| 18 | Gemini 3.5 Flash Google DeepMind | 1475.5 1470.9–1480.2 |
| 19 | Gemini 3.5 Flashmedium effort Google DeepMind | 1474.9 1468.1–1481.7 |
| 20 | Grok 4.20 Beta1 xAI | 1474.6 1469.9–1479.2 |
| 21 | GPT 5.5 Instant OpenAI | 1473.7 1468.6–1478.8 |
| 22 | Qwen3 7 Max Preview Alibaba | 1473.6 1463.6–1483.6 |
| 23 | Claude Opus 4.8 Anthropic | 1472.8 1468.6–1477.1 |
| 24 | Claude Opus 4.5 20251101 Thinking 32K Anthropic | 1472.8 1468.9–1476.7 |
| 25 | Claude Sonnet 4.6 Anthropic | 1472.4 1468.8–1476.0 |
One row per evaluated system — reasoning-effort variants rank separately. Pool-relative ratings (Elo) are only meaningful within this single arena — never compare them across sources.
