ELO Rating78 models ranked

Coding ELO Leaderboard 2026

Coding ELO is a specialised Chatbot Arena leaderboard restricted to coding tasks. It measures human preference when comparing model-generated code, debugging explanations, and programming help.

Quick Answer

The best model on Coding ELO in 2026 is Claude Opus 4 by Anthropic, scoring 1497 ELO. Runner-up: Claude Sonnet 4 (1473).

78 / 78 models
#ModelScore
🥇Claude Opus 41497ELO
🥈Claude Sonnet 41473ELO
🥉Gemini 2.5 Pro1445ELO
4DeepSeek R11398ELO
5Qwen 2.5 Max1374ELO
6Gemini 2.0 Flash1360ELO
7DeepSeek V31358ELO
8o3-mini1348ELO
9o31320ELO
10o11310ELO
11Claude 3.5 Sonnet1300ELO
12GPT-4.51295ELO
13Qwen 3 235B MoE1290ELO
14Qwen 2.5 Coder 32B1290ELO
15Gemini Experimental 12061280ELO
16Llama 4 Maverick1280ELO
17Grok 31280ELO
18o1-mini1280ELO
19DeepSeek R1 (Groq)1270ELO
20DeepSeek R1 (Together)1270ELO
21Gemini 2.0 Flash Thinking1270ELO
22ChatGPT-4o Latest1270ELO
23o4-mini1270ELO
24GPT-4o (Aug 2024)1265ELO
25Gemini 2.5 Flash1260ELO
26QwQ 32B1250ELO
27Claude 3.5 Haiku1250ELO
28Codestral 22B1250ELO
29GPT-4 Turbo1245ELO
30DeepSeek R1 Distill Llama 70B1240ELO
31Mistral Large1240ELO
32DeepSeek R1 Distill Qwen 32B1235ELO
33Sonar Reasoning1235ELO
34Llama 4 Scout1230ELO
35Command A1230ELO
36Grok 21225ELO
37DeepSeek V2.51220ELO
38Mixtral 8x22B (Fireworks)1220ELO
39WizardLM-2 8x22B1220ELO
40Qwen 2.5 72B1210ELO
41Qwen 2.5 72B (Together)1210ELO
42Amazon Nova Pro1210ELO
43GPT-4 11210ELO
44Llama 3.1 405B (Fireworks)1200ELO
45Llama 3.1 405B1200ELO
46Llama 3.1 405B (Together)1200ELO
47GPT-4 1.5-mini1200ELO
48Claude Haiku 41195ELO
49Sonar Pro1195ELO
50Gemma 2 27B1195ELO
51Yi-Large1195ELO
52Phi-3.5 MoE1190ELO
53Grok 3-mini1190ELO
54Llama 3.3 70B (Fireworks)1180ELO
55Llama 3.3 70B (Groq)1180ELO
56Llama 3.3 70B1180ELO
57Llama 3.3 70B (Together)1180ELO
58Phi-3 Medium1175ELO
59Gemini 2.0 Flash Lite1170ELO
60Yi-Lightning1170ELO
61Sonar1170ELO
62Phi-3.5 Mini1165ELO
63GPT-4o1162ELO
64Command R+1160ELO
65Mistral Small1160ELO
66Amazon Nova Lite1160ELO
67Gemma 2 9B (Groq)1160ELO
68GPT-4 1.5-nano1160ELO
69Gemma 2 9B1155ELO
70Mixtral 8x7B (Groq)1150ELO
71InternLM 2.5 20B1150ELO
72Phi-41130ELO
73Command R1100ELO
74Amazon Nova Micro1100ELO
75Command R7B1100ELO
76GPT-3.5 Turbo1100ELO
77GPT-4o Mini1098ELO
78Mistral 7B (Together)1090ELO

What Coding ELO Tests

Human preference on coding-specific tasks: writing functions, explaining bugs, reviewing pull requests. Derived from the same Arena methodology applied only to code conversations.

Score Range

1100–1450+ (average ~1220)

Compare models side-by-side

Full spec comparison — pricing, context window, and all benchmarks.

Compare Models →