Skip to content

AI Benchmark Dashboard

Compare 41+ AI models across 8 standardized benchmarks. Scores sourced from Papers with Code, Hugging Face Open LLM Leaderboard, and official model cards. Updated regularly with the latest results.

8 benchmarks
41 models tracked
4 categories
o32 wins
Claude Opus 4.6 (thinking)1 win
Qwen3.7 Max1 win
#4Claude Sonnet 4.51 win

LMSYS Chatbot Arena Elo

Crowdsourced blind human preference rankings. Users compare anonymous model outputs — the most democratic LLM benchmark.

reasoning11 modelsUpdated July 27, 2026
Source
Claude Opus 4.6 (thinking)Anthropic
1501.0Elo
GPT-4oOpenAI
1287.0Elo
Claude Opus 4.6Anthropic
1280.0Elo
4
Gemini 2.0 UltraGoogle
1275.0Elo
5
Grok 3xAI
1268.0Elo
6
Claude Sonnet 4.6Anthropic
1260.0Elo
7
DeepSeek V3DeepSeek
1248.0Elo
8
Llama 4 MaverickMeta
1240.0Elo
9
Qwen 3 235BAlibaba
1232.0Elo
10
GPT-4o MiniOpenAI
1200.0Elo
11
Mistral Large 2.5Mistral
1195.0Elo

MMLU

Massive Multitask Language Understanding — measures knowledge across 57 academic subjects including STEM, humanities, and social sciences.

knowledge17 modelsUpdated July 27, 2026
Source
Qwen3.7 MaxAlibaba
93.7%
GPT-5OpenAI
93.5%
o3OpenAI
93.1%
4
GPT-5.5OpenAI
92.4%
5
Claude Opus 4.6Anthropic
92.3%
6
Gemini 2.0 UltraGoogle
90.8%
7
DeepSeek R1DeepSeek
90.8%
8
Claude Sonnet 4.6Anthropic
89.5%
9
GPT-4oOpenAI
88.7%
10
Llama 4 MaverickMeta
88.2%
11
DeepSeek V3DeepSeek
87.1%
12
Qwen 3 235BAlibaba
86.5%
13
Mistral Large 2.5Mistral
86.3%
14
Grok 3xAI
85.7%
15
Gemini 2.0 FlashGoogle
83.4%
16
GPT-4o MiniOpenAI
82.0%
17
Claude Haiku 4.5Anthropic
80.1%

HumanEval

Evaluates code generation by asking models to complete Python functions from docstrings. Tests practical programming ability.

coding14 modelsUpdated July 27, 2026
Source
Claude Sonnet 4.5Anthropic
97.6%
Grok 4.1 FastxAI
97.0%
Claude Opus 4.6Anthropic
93.7%
4
GPT-4oOpenAI
90.2%
5
DeepSeek V3DeepSeek
89.8%
6
Gemini 2.0 UltraGoogle
88.4%
7
Claude Sonnet 4.6Anthropic
88.0%
8
Qwen 3 235BAlibaba
87.5%
9
Llama 4 MaverickMeta
85.5%
10
Grok 3xAI
85.0%
11
Mistral Large 2.5Mistral
84.1%
12
GPT-4o MiniOpenAI
80.5%
13
Gemini 2.0 FlashGoogle
79.2%
14
Claude Haiku 4.5Anthropic
78.3%

ARC-AGI

Abstraction and Reasoning Corpus — tests novel pattern recognition and abstract reasoning. Considered the hardest AI benchmark.

reasoning12 modelsUpdated July 27, 2026
Source
GPT-5.6 Sol (ARC-AGI-2)OpenAI
92.5%
Claude Opus 5 (ARC-AGI-2)Anthropic
90.4%
o3 (high compute)OpenAI
87.5%
4
GPT-5.5 (ARC-AGI-2)OpenAI
85.0%
5
Gemini 3 Deep Think (ARC-AGI-2)Google
84.6%
6
GPT-5.4 Pro (ARC-AGI-2)OpenAI
83.3%
7
Gemini 3.1 Pro (ARC-AGI-2)Google DeepMind
77.1%
8
o3-miniOpenAI
77.0%
9
Claude Opus 4.6Anthropic
53.0%
10
DeepSeek R1DeepSeek
42.0%
11
Gemini 2.0 UltraGoogle
38.5%
12
GPT-4oOpenAI
21.0%

GPQA Diamond

Graduate-level science questions written by PhD experts. Extremely hard — even domain experts only score ~65%.

reasoning17 modelsUpdated July 27, 2026
Source
Claude Fable 5Anthropic
94.6%
Claude Opus 4.7Anthropic
94.2%
Gemini 3.1 ProGoogle DeepMind
94.1%
4
GPT-5.6 SolOpenAI
94.1%
5
Claude Opus 4.8Anthropic
93.6%
6
Kimi K2.6Moonshot AI
90.8%
7
DeepSeek V4 Pro-MaxDeepSeek
90.1%
8
o3OpenAI
87.7%
9
GPT-5.5OpenAI
85.7%
10
o1OpenAI
78.0%
11
Claude Opus 4.6Anthropic
74.9%
12
Gemini 2.0 UltraGoogle
72.1%
13
DeepSeek R1DeepSeek
71.5%
14
Claude Sonnet 4.6Anthropic
65.0%
15
Llama 4 MaverickMeta
58.3%
16
Grok 3xAI
56.7%
17
GPT-4oOpenAI
53.6%

SWE-bench Verified

Tests ability to solve real GitHub issues from popular Python repos. The gold standard for agentic coding evaluation.

coding16 modelsUpdated July 27, 2026
Source
GPT-5.6 SolOpenAI
96.2%
Claude Opus 5Anthropic
96.0%
Claude Fable 5Anthropic
95.5%
4
Claude Mythos 5Anthropic
95.5%
5
GPT-5.5OpenAI
88.7%
6
Claude Opus 4.8Anthropic
88.6%
7
Claude Opus 4.7Anthropic
87.6%
8
Gemini 3.1 ProGoogle DeepMind
80.6%
9
Claude Opus 4.6 + ScaffoldAnthropic
72.0%
10
o3 + ScaffoldOpenAI
69.1%
11
Claude Sonnet 4.6Anthropic
55.3%
12
GPT-4.1OpenAI
54.6%
13
DeepSeek R1DeepSeek
49.2%
14
Gemini 2.0 UltraGoogle
47.8%
15
GPT-4oOpenAI
38.4%
16
Llama 4 MaverickMeta
32.1%

MATH

Competition-level math problems from AMC, AIME, and Olympiad. Tests mathematical reasoning and problem-solving.

math11 modelsUpdated March 20, 2026
Source
o3OpenAI
96.7%
o1OpenAI
94.8%
Claude Opus 4.6Anthropic
85.1%
4
Gemini 2.0 UltraGoogle
83.9%
5
DeepSeek R1DeepSeek
83.5%
6
Qwen 3 235BAlibaba
81.2%
7
Llama 4 MaverickMeta
79.1%
8
GPT-4oOpenAI
76.6%
9
Grok 3xAI
76.0%
10
Mistral Large 2.5Mistral
74.5%
11
Claude Sonnet 4.6Anthropic
73.8%

GSM8K

Grade School Math — 8,500 grade school math word problems. Tests basic mathematical reasoning and multi-step problem solving.

math10 modelsUpdated March 18, 2026
Source
o3OpenAI
99.2%
Claude Opus 4.6Anthropic
97.8%
Gemini 2.0 UltraGoogle
96.1%
4
GPT-4oOpenAI
95.8%
5
DeepSeek V3DeepSeek
95.0%
6
Claude Sonnet 4.6Anthropic
94.5%
7
Llama 4 MaverickMeta
93.7%
8
Qwen 3 235BAlibaba
93.2%
9
GPT-4o MiniOpenAI
87.0%
10
Gemini 2.0 FlashGoogle
86.5%

Frequently Asked Questions

What is the best AI model in 2026?

It depends on the task. Qwen3.7 Max leads MMLU (93.7%); Claude Sonnet 4.5 tops HumanEval (97.6%); GPT-5.6 Sol leads SWE-bench (96.2%); Claude Fable 5 leads GPQA Diamond (94.6%) — based on the latest official scores in our dashboard.

What is the MMLU benchmark?

MMLU (Massive Multitask Language Understanding) tests AI models across 57 subjects including STEM, humanities, and social sciences. It measures broad knowledge and reasoning ability. Scores above 90% indicate expert-level performance.

What is HumanEval?

HumanEval measures AI code generation using 164 hand-written Python programming problems. Models must generate correct, functional code that passes test cases. Scores above 90% indicate near-human coding ability.

How are AI benchmarks scored?

Most benchmarks use accuracy percentage (0-100%). Some use Elo ratings (Chatbot Arena) or pass rates (SWE-bench). Scores come from official evaluations by model providers and independent organizations.

Data sourced from Papers with Code, Hugging Face, and official model evaluations. Updated regularly. Scores reflect official reported results only.