AI & Tech Intelligence · Epoch AI
AI Model Benchmarks: The Leaderboard
A source-cited AI model benchmark leaderboard — 18 frontier models scored on GPQA Diamond, SWE-Bench Verified, FrontierMath, MATH and more, from independent Epoch AI evaluations. Which AI is best at science, code and maths. CC BY 4.0.
TL;DR — Kimi K3 leads GPQA Diamond (graduate-level science) at 93.1%.
Epoch AI · Epoch AI | 18 entities | 854 words · 13 sections | data: CSV + JSON
Executive summary
This is a source-cited leaderboard of the leading AI models, scored on the benchmarks that matter: graduate-level science (GPQA Diamond), real-world software engineering (SWE-Bench Verified), research-level mathematics (FrontierMath), the hardest competition maths (MATH Level 5) and factual accuracy (SimpleQA). All 18 models are ranked on the same independent evaluations, run by Epoch AI rather than self-reported by the labs. The headline: there is no single 'best' model — Kimi K3 tops science reasoning, GPT-5 leads coding, and every model still struggles on FrontierMath. Each benchmark is charted and defined below; the full table and data are downloadable under CC BY 4.0.
“These are extremely challenging. I think that in the near term basically the only way to solve them, short of having a real domain expert in the area, is by a combination of a semi-expert like a graduate student in a related field, maybe paired with some combination of a modern AI and lots of other algebra packages.”
Key findings
Best on graduate-level science
Kimi K3 (Moonshot) leads GPQA Diamond at 93.1% — a test of graduate-level science reasoning where human experts score around 65%.
Source: Epoch AI Benchmarking Hub · 2026 · confidence: High
Best at real-world coding
GPT-5 leads SWE-Bench Verified at 73.6%, solving real GitHub software-engineering tasks — the benchmark closest to what developers actually do.
Source: Epoch AI Benchmarking Hub · 2026 · confidence: High
FrontierMath remains unsolved
Even the best model scores only 32.4% on FrontierMath, a set of research-level maths problems built to resist AI — evidence that hard mathematical reasoning is far from solved.
Source: Epoch AI Benchmarking Hub · 2026 · confidence: High
Data vintage — August 2026
Scores from Epoch AI's Benchmarking Hub — independent evaluations (best across scorers) on public and private benchmark sets. Updated as new models are released and evaluated; this leaderboard is a snapshot.
The leaderboard — every model, every benchmark
All 18 benchmarked models, ranked by GPQA Diamond (graduate-level science). Scores are the best across scorers from independent Epoch AI evaluations; '—' means the model was not run on that benchmark.
| # | Model | Developer | GPQA Diamond | SWE-Bench Verified | FrontierMath | MATH Level 5 | SimpleQA | AIME (OTIS Mock) |
|---|---|---|---|---|---|---|---|---|
| 1 | Kimi K3 | Moonshot | 93.1% | — | — | — | 42.7% | 97.2% |
| 2 | Grok 4 | xAI | 87.0% | — | 19.7% | — | 47.9% | 84.0% |
| 3 | GPT-5 | OpenAI | 86.2% | 73.6% | 32.4% | 98.1% | 50.6% | 91.4% |
| 4 | Claude 3.7 Sonnet | Anthropic | 79.7% | 61.0% | 4.1% | 91.2% | — | 57.8% |
| 5 | Grok 3 | xAI | 75.8% | — | 3.8% | 88.7% | — | 55.6% |
| 6 | Qwen3-Max | Alibaba | 72.6% | — | — | 97.1% | 67.5% | 73.3% |
| 7 | GPT-4.5 | OpenAI | 68.7% | — | — | 78.6% | — | 37.8% |
| 8 | Gemini 1.5 Pro | Google DeepMind | 57.2% | — | — | 70.4% | — | 23.1% |
| 9 | Claude 3.5 Sonnet | Anthropic | 54.0% | — | 1.0% | 51.7% | — | 6.5% |
| 10 | Grok-2 | xAI | 53.8% | — | 0.7% | 63.5% | — | 11.5% |
| 11 | Llama 3.1-405B | Meta AI | 50.9% | — | — | 49.8% | — | 9.7% |
| 12 | GPT-4o | OpenAI | 49.2% | 31.0% | 0.3% | 53.3% | — | 6.4% |
| 13 | Claude 3 Opus | Anthropic | 47.2% | — | — | 37.5% | — | 4.7% |
| 14 | GPT-4 Turbo (Apr 2024) | OpenAI | 46.6% | — | — | 46.7% | — | 6.7% |
| 15 | GPT-4 Turbo (Nov 2023) | OpenAI | 42.4% | — | — | 40.0% | — | — |
| 16 | GPT-4 (Mar 2023) | OpenAI | 35.7% | — | — | — | — | 0.6% |
| 17 | Claude 2 | Anthropic | 34.7% | — | — | 11.7% | — | 2.5% |
| 18 | GPT-4 (Jun 2023) | OpenAI | 30.7% | — | — | 23.0% | — | 1.1% |
What the benchmarks reveal
The frontier is crowded at the top. On GPQA Diamond — graduate-level science questions experts get ~65% on — the leader is Kimi K3 (Moonshot) at 93.1%, but several models from different labs sit within a few points. No single lab owns every benchmark: capability leadership is split across OpenAI, Anthropic, Google, xAI and fast-rising Chinese labs.
Different benchmarks crown different winners. GPT-5 leads real-world software engineering (SWE-Bench Verified, 73.6%) while GPT-5 tops the hardest competition maths (MATH Level 5, 98.1%). A model that dominates one skill can trail on another — which is why a single 'best AI model' number is misleading.
FrontierMath is the benchmark that still humbles every model. Designed with research mathematicians to be extremely hard, the best score here is just 32.4% (GPT-5) — a reminder that frontier maths reasoning remains far from solved even as models saturate easier tests.
These are independent evaluations run by Epoch AI, not lab-reported figures — the cleanest way to compare models on a level field. Scores are the best across scorers and move as new models ship; the leaderboard is a snapshot, not a verdict.
Who leads GPQA Diamond?
Kimi K3 (Moonshot) leads GPQA Diamond at 93.1%. Graduate-level, 'Google-proof' science questions (biology, physics, chemistry); human experts score ~65%.
Who leads SWE-Bench Verified?
GPT-5 (OpenAI) leads SWE-Bench Verified at 73.6%. Real GitHub software-engineering tasks — the benchmark closest to practical coding work.
Who leads FrontierMath?
GPT-5 (OpenAI) leads FrontierMath at 32.4%. Research-level mathematics problems designed with mathematicians to be extremely hard for AI.
Who leads MATH Level 5?
GPT-5 (OpenAI) leads MATH Level 5 at 98.1%. The hardest tier of competition mathematics problems.
How to read AI benchmarks
AI benchmarks measure narrow, specific skills — not general intelligence. A high GPQA score means strong science-question answering; it says little about coding or factual reliability. Benchmarks also saturate: once models cluster near 100%, the test stops discriminating, which is why harder sets like FrontierMath and SWE-Bench matter most at the frontier.
Independent evaluation matters. Lab-reported scores use different prompts, tools and scoring, making cross-model comparison unreliable. Epoch AI runs models on the same harness, which is why this leaderboard uses its numbers. Even so, treat any single score as one data point, not a verdict on which model is 'best'.
Scoreboard (machine-readable data)
Every headline indicator with its value, period, source and confidence. Free to reuse under CC BY 4.0.
| Indicator | Value | Period | Source | Conf. |
|---|---|---|---|---|
| Kimi K3 | 93.1 pct | 2026 | Epoch AI Benchmarking Hub | High |
| Grok 4 | 87 pct | 2026 | Epoch AI Benchmarking Hub | High |
| GPT-5 | 86.2 pct | 2026 | Epoch AI Benchmarking Hub | High |
| Claude 3.7 Sonnet | 79.7 pct | 2026 | Epoch AI Benchmarking Hub | High |
| Grok 3 | 75.8 pct | 2026 | Epoch AI Benchmarking Hub | High |
| Qwen3-Max | 72.6 pct | 2026 | Epoch AI Benchmarking Hub | High |
| GPT-4.5 | 68.7 pct | 2026 | Epoch AI Benchmarking Hub | High |
| Gemini 1.5 Pro | 57.2 pct | 2026 | Epoch AI Benchmarking Hub | High |
| Claude 3.5 Sonnet | 54 pct | 2026 | Epoch AI Benchmarking Hub | High |
| Grok-2 | 53.8 pct | 2026 | Epoch AI Benchmarking Hub | High |
| Llama 3.1-405B | 50.9 pct | 2026 | Epoch AI Benchmarking Hub | High |
| GPT-4o | 49.2 pct | 2026 | Epoch AI Benchmarking Hub | High |
| Claude 3 Opus | 47.2 pct | 2026 | Epoch AI Benchmarking Hub | High |
| GPT-4 Turbo (Apr 2024) | 46.6 pct | 2026 | Epoch AI Benchmarking Hub | High |
| GPT-4 Turbo (Nov 2023) | 42.4 pct | 2026 | Epoch AI Benchmarking Hub | High |
| GPT-4 (Mar 2023) | 35.7 pct | 2026 | Epoch AI Benchmarking Hub | High |
| Claude 2 | 34.7 pct | 2026 | Epoch AI Benchmarking Hub | High |
| GPT-4 (Jun 2023) | 30.7 pct | 2026 | Epoch AI Benchmarking Hub | High |
Methodology & verification
Scores from Epoch AI's Benchmarking Hub (CC BY 4.0): independent evaluations of AI models on standardised benchmark sets, reporting the best score across scorers. Benchmarks: GPQA Diamond (graduate-level science Q&A), SWE-Bench Verified (real GitHub software-engineering tasks), FrontierMath (research-level mathematics), MATH Level 5 (hardest competition maths), AIME (olympiad maths) and SimpleQA (factual accuracy). Models are ranked by GPQA Diamond where available. Scores are point-in-time and move as models are updated. Nothing is estimated by Affärslivet. Confidence: High.
Data dictionary
| Field | Type | Description |
|---|---|---|
| model | string | AI model / version |
| value | number | GPQA Diamond score, % (Epoch AI) |
| geography | string | Developer organisation |
Frequently asked questions
What is the best AI model in 2026?
There is no single best model. On graduate-level science (GPQA Diamond) Kimi K3 leads at 93.1%; on real-world coding (SWE-Bench Verified) GPT-5 leads at 73.6%. Leadership is split across benchmarks and labs. Source: Epoch AI.
Which AI model is best at coding?
On SWE-Bench Verified — real GitHub software-engineering tasks — GPT-5 leads at 73.6%. It is the benchmark closest to what developers actually do. Source: Epoch AI.
What is GPQA Diamond?
GPQA Diamond is a benchmark of graduate-level, 'Google-proof' science questions in biology, physics and chemistry. Human experts score around 65%; the leading AI model reaches 93.1%. Source: Epoch AI.
What is the hardest AI benchmark?
Among common benchmarks, FrontierMath is the hardest — research-level maths problems designed with mathematicians to resist AI. The best model scores only 32.4%, versus near-saturation on easier maths tests. Source: Epoch AI.
Are these scores reported by the AI labs?
No. These are independent evaluations run by Epoch AI on a common harness, not self-reported figures. That makes cross-model comparison far more reliable, since lab-reported scores use different prompts, tools and scoring.
Do AI benchmarks measure intelligence?
No. Benchmarks measure narrow, specific skills — science Q&A, coding, maths, factual recall — not general intelligence. A model can top one benchmark and trail on another, so no single score captures overall capability.
Which countries rank in the top 5 of the AI Model Benchmarks: The Leaderboard?
The top five are Kimi K3, Grok 4, GPT-5, Claude 3.7 Sonnet and Grok 3, based on the AI Model Benchmarks: The Leaderboard (Affärslivet, 2026-08-17).
How many countries are ranked in the AI Model Benchmarks: The Leaderboard?
18 economies are included in the AI Model Benchmarks: The Leaderboard (2026-08-17). Full data is free to download as CSV or JSON under CC BY 4.0.
Glossary
- GPQA Diamond
- Graduate-level, 'Google-proof' science questions (biology, physics, chemistry); human experts score ~65%. ↗
- SWE-Bench Verified
- Real GitHub software-engineering tasks — the benchmark closest to practical coding work. ↗
- FrontierMath
- Research-level mathematics problems designed with mathematicians to be extremely hard for AI. ↗
- MATH Level 5
- The hardest tier of competition mathematics problems. ↗
- SimpleQA
- Short factual questions measuring accuracy and resistance to hallucination. ↗
- AIME (OTIS Mock)
- Olympiad-level mathematics (American Invitational Mathematics Examination style). ↗
Embed & cite this report
Free to reuse under CC BY 4.0. Embed the live-updating widget on your site, or cite the report directly — always with attribution to Affärslivet.
Embed (HTML) — auto-updating
<iframe src="https://xn--affrslivet-s5a.com/en/embed/ai-model-benchmarks" width="100%" height="520" style="border:1px solid #e3e3e6" title="AI Model Benchmarks: The Leaderboard — Affärslivet" loading="lazy"></iframe> <p style="font:12px sans-serif">Source: <a href="https://xn--affrslivet-s5a.com/en/reports/ai-model-benchmarks">Affärslivet</a></p>
APA
Affärslivet Research. (2026). AI Model Benchmarks: The Leaderboard. Affärslivet. Version 1.0. https://xn--affrslivet-s5a.com/en/reports/ai-model-benchmarks
MLA
Affärslivet Research. "AI Model Benchmarks: The Leaderboard." Affärslivet, 2026-07-30, https://xn--affrslivet-s5a.com/en/reports/ai-model-benchmarks.
BibTeX
@techreport{affarslivet_ai_model_benchmarks,
title = {AI Model Benchmarks: The Leaderboard},
author = {{Affärslivet Research}},
year = {2026},
note = {Version 1.0},
url = {https://xn--affrslivet-s5a.com/en/reports/ai-model-benchmarks}
} License CC BY 4.0 — free to cite, embed and republish with attribution to Affärslivet. Data also as CSV / JSON.
Sources
Part of Affärslivet AI Intelligence
This report is one part of Affärslivet's source-cited AI knowledge layer. Start with the big picture:
- The State of AI 2026 — our free, source-cited annual report
- AI Statistics 2026 — 230+ source-cited AI facts & figures
- The AI Intelligence hub — all our indexes, rankings, model & lab profiles