ELO Rating90 models ranked

Reasoning ELO Leaderboard 2026

Reasoning ELO is the Chatbot Arena leaderboard filtered to hard reasoning and math problems. It measures how well models solve multi-step logic, quantitative reasoning, and complex problem-solving.

Quick Answer

The best model on Reasoning ELO in 2026 is Claude Opus 4 by Anthropic, scoring 1497 ELO. Runner-up: Claude Sonnet 4 (1473).

90 / 90 models
#ModelScore
🥇Claude Opus 41497ELO
🥈Claude Sonnet 41473ELO
🥉Gemini 2.5 Pro1445ELO
4DeepSeek R11398ELO
5Qwen 2.5 Max1374ELO
6Gemini 2.0 Flash1360ELO
7DeepSeek V31358ELO
8o31350ELO
9o3-mini1348ELO
10o11330ELO
11Qwen 3 235B MoE1320ELO
12Gemini Experimental 12061310ELO
13DeepSeek R1 (Groq)1300ELO
14DeepSeek R1 (Together)1300ELO
15Grok 31295ELO
16GPT-4.51290ELO
17Gemini 2.0 Flash Thinking1290ELO
18o1-mini1280ELO
19Llama 4 Maverick1275ELO
20o4-mini1275ELO
21Claude 3.5 Sonnet1270ELO
22Gemini 2.5 Flash1270ELO
23QwQ 32B1270ELO
24ChatGPT-4o Latest1265ELO
25DeepSeek R1 Distill Llama 70B1260ELO
26GPT-4o (Aug 2024)1255ELO
27DeepSeek R1 Distill Qwen 32B1250ELO
28Llama 3.1 405B (Fireworks)1250ELO
29Llama 3.1 405B1250ELO
30Sonar Reasoning1250ELO
31Llama 3.1 405B (Together)1250ELO
32Command A1240ELO
33GPT-4 Turbo1240ELO
34Grok 21240ELO
35Mistral Large1230ELO
36Gemini 1.5 Pro1230ELO
37Grok 2 Vision1230ELO
38Pixtral Large1230ELO
39Qwen 2.5 72B1230ELO
40Qwen 2.5 72B (Together)1230ELO
41Llama 4 Scout1220ELO
42Amazon Nova Pro1220ELO
43Claude 3.5 Haiku1220ELO
44Llama 3.3 70B (Fireworks)1220ELO
45Llama 3.3 70B (Groq)1220ELO
46Llama 3.3 70B1220ELO
47Mistral Medium 31220ELO
48Llama 3.3 70B (Together)1220ELO
49Llama 3.2 90B Vision1210ELO
50Sonar Pro1210ELO
51DeepSeek V2.51200ELO
52Mixtral 8x22B (Fireworks)1200ELO
53GPT-4 11200ELO
54WizardLM-2 8x22B1200ELO
55Llama 3.1 70B1195ELO
56Phi-3.5 MoE1195ELO
57Gemini 1.5 Flash1190ELO
58Gemma 2 27B1190ELO
59Claude Haiku 41185ELO
60Yi-Large1185ELO
61GPT-4 1.5-mini1180ELO
62Grok 3-mini1175ELO
63Command R+1170ELO
64Amazon Nova Lite1170ELO
65Gemma 2 9B (Groq)1170ELO
66Phi-3 Medium1170ELO
67Yi-Lightning1165ELO
68GPT-4o1162ELO
69Gemini 2.0 Flash Lite1160ELO
70Gemma 2 9B1160ELO
71Mixtral 8x7B (Groq)1160ELO
72Llama 3.2 11B Vision1160ELO
73Phi-3.5 Mini1160ELO
74Qwen 2.5 7B1160ELO
75Sonar1160ELO
76InternLM 2.5 20B1155ELO
77Mistral Small1150ELO
78Gemini 1.5 Flash 8B1150ELO
79GPT-4 1.5-nano1150ELO
80Phi-41140ELO
81Mistral Nemo 12B1140ELO
82Amazon Nova Micro1130ELO
83Command R7B1120ELO
84GPT-3.5 Turbo1120ELO
85Llama 3.1 8B (Groq)1120ELO
86Llama 3.1 8B1120ELO
87Command R1110ELO
88Mistral 7B1100ELO
89Mistral 7B (Together)1100ELO
90GPT-4o Mini1098ELO

What Reasoning ELO Tests

Human preference on reasoning-heavy tasks: math word problems, logic puzzles, structured analysis. A higher score means humans find the model's reasoning more sound and useful.

Score Range

1100–1450+ (average ~1210)

Compare models side-by-side

Full spec comparison — pricing, context window, and all benchmarks.

Compare Models →