Benchmark evidence
AI benchmark leaderboards
Browse benchmark results by the organisations and research labs that publish them: Cursor, LMSYS Arena, Artificial Analysis, DeepSWE, SWE-bench Verified, LiveBench, and Provider Catalogs. Compare Pareto efficiency (accuracy vs cost/tokens/steps) and raw score tables.
Catalogue review recorded · individual claims have separate datesStructured datasetSourceMethodology
Swipe sideways to see all publishers.
Highest displayed publisher result
Claude Fable 5.1
66 (Index score)
Rows shown
25 rows
Index score
Snapshot date
2026-09-08
Cohort edition v4.1.1 era — historical rows stay under their captured edition
Artificial Analysis Intelligence Index ↗
Understanding Intelligence Index
A composite index aggregating several evaluations into a comparative model-quality signal.
What it reveals
It is useful as one broad reference point when a buyer needs an initial shortlist.
What it cannot establish
Composite scores hide task-specific variance, version changes, and the trade-off between quality and cost.
These are transcribed publisher results, not tests run by AI Bench Index. A publisher’s order may include different model settings. Row sources and notes appear in the full table; the snapshot date is not a test-run date.
Read the publisher’s methodology and results ↗Score profile
Index score by published row
Higher is better
Artificial Analysis composite intelligence index across leading frontier models.
Catalogued suites
Standardized evaluation suites
Dataset definitions, contamination concerns, and recorded evaluation methods. Historical suites remain available as reference.
MMLU-Pro
Massive Multitask Language Understanding Professional: 12,000+ complex reasoning questions across 14 disciplines with 10 answer options, which reduce (but do not eliminate) the chance of a correct random guess.
- Measures
- Hard multiple-choice questions across 14 academic subjects with 10 options.
- Does not test
- Tool use, long-horizon coding, or real repository patches.
- Misread
- Not interchangeable with original 5-option MMLU.
No pooled distribution: recorded evaluation conditions differ or are missing.
Comparability notes
Dataset 2024-06-01. Contamination low. Scores are never blended across harness versions.
SWE-bench Verified
Human-validated subset of 500 real GitHub issue resolution tasks evaluating end-to-end software engineering on production codebases.
- Measures
- Resolving real GitHub issues on a 500-task human-validated subset.
- Does not test
- Chat coding quality in an IDE, or SWE-bench Pro / multilingual variants.
- Misread
- Vals mini-swe-agent 2026 rows are not the same experiment as 2024 two-tool or Aider scaffolds.
Comparability notes
Dataset 2024-08-15. Contamination high. Scores are never blended across harness versions.
GPQA Diamond
Google-Proof Q&A: 198 graduate-level biology, chemistry, and physics questions authored and validated by PhD domain experts.
- Measures
- 198 graduate-level biology, chemistry, and physics questions written by domain experts.
- Does not test
- Laboratory skill, tool use, or coding.
- Misread
- Diamond is a subset of GPQA; do not mix Diamond with the full set.
No pooled distribution: recorded evaluation conditions differ or are missing.
Comparability notes
Dataset 2023-11-20. Contamination low. Scores are never blended across harness versions.
MATH 500
superseded suite · retained as historical reference, not an active frontier ranking.
Representative 500-problem subset of challenging competition-level high school mathematics from the Hendrycks MATH benchmark.
- Measures
- A 500-problem slice of Hendrycks MATH competition items.
- Does not test
- Proof writing or MATH-500 replacements such as AIME live contests.
- Misread
- This suite is catalogued as superseded for active leaderboards.
No pooled distribution: recorded evaluation conditions differ or are missing.
Comparability notes
Dataset 2024-05-01. Contamination moderate. Scores are never blended across harness versions.
AIME 2024
American Invitational Mathematics Examination 2024: 30 non-multiple choice competition questions for the top 5% of math students.
- Measures
- AIME 2024 contest problems (integer answers).
- Does not test
- Open-ended proof contests or MATH-500.
- Misread
- AIME 2024 is not AIME 2025.
No pooled distribution: recorded evaluation conditions differ or are missing.
Comparability notes
Dataset 2024-02-15. Contamination low. Scores are never blended across harness versions.
LiveBench
Contamination-resistant benchmark with questions sourced from post-cutoff competitions, arXiv papers, and real news. LiveBench was introduced with monthly question releases; the current public site describes a six-month refresh cadence. Results are only compared within the same release and configuration.
- Measures
- A frequently refreshed mix of tasks intended to reduce contamination.
- Does not test
- A single skill such as only coding or only math.
- Misread
- Snapshots from different dates are not one ladder.
No pooled distribution: recorded evaluation conditions differ or are missing.
Comparability notes
Dataset 2024-10-01. Contamination low. Scores are never blended across harness versions.
HumanEval
164 hand-crafted Python programming problems with unit tests measuring functional correctness (pass@1) from docstrings.
- Measures
- Short Python function synthesis from docstrings (pass@k).
- Does not test
- Multi-file repository repair (that is SWE-bench).
- Misread
- HumanEval is not SWE-bench Verified.
No pooled distribution: recorded evaluation conditions differ or are missing.
Comparability notes
Dataset 2021-07-07. Contamination high. Scores are never blended across harness versions.
Arena-Hard Auto
500 challenging user queries from Chatbot Arena evaluated with automated LLM judges, showing 89.1% correlation with crowdsourced human votes.
- Measures
- Hard conversational prompts judged against a baseline policy.
- Does not test
- Deterministic unit-tested coding.
- Misread
- LMSYS Arena Elo boards are separate from Arena-Hard.
Median 85.2 across 5 models with evidence (range 79.3–86.8).
Comparability notes
Dataset 2024-06-15. Contamination low. Scores are never blended across harness versions.
BigCodeBench
1,140 challenging Python programming tasks invoking 139 external libraries, evaluating real-world software development.
- Measures
- More realistic Python programming tasks than HumanEval.
- Does not test
- GitHub issue resolution on existing repos.
- Misread
- Instruct vs complete splits are not interchangeable.
Comparability notes
Dataset 2024-06-25. Contamination low. Scores are never blended across harness versions.
SWE-bench Lite
300 issue resolution tasks selected for clean self-contained evaluation across 12 major Python repositories.
- Measures
- A smaller SWE-bench subset than Verified.
- Does not test
- The 500-task Verified set.
- Misread
- Lite is not Verified.
Comparability notes
Dataset 2024-03-01. Contamination low. Scores are never blended across harness versions.
Coverage matrix
Which tracked models have a non-superseded receipt on each active suite. Missing cells are missing evidence, not zeros.
| Model | MMLU-Pro | SWE-bench Verified | GPQA Diamond | AIME 2024 | LiveBench | HumanEval | Arena-Hard Auto | BigCodeBench | SWE-bench Lite |
|---|---|---|---|---|---|---|---|---|---|
| OpenAI o1 | — | — | 75.7 | 79.2 | 68.2 | — | 86.8 | — | — |
| OpenAI o1-mini | — | — | — | — | — | — | — | — | — |
| GPT-4o | 72.6 | — | 50.3 | 9.3 | 59.2 | 90.2 | 79.3 | — | 38 |
| GPT-4o mini | 62.5 | — | 43.4 | — | — | — | — | — | — |
| Claude 3.5 Sonnet | 78 | — | 65 | 16 | 64.8 | 93.7 | 85.2 | — | 43.3 |
| Claude 3.5 Haiku | — | — | — | — | — | — | — | — | — |
| Claude 3 Opus | — | — | — | — | — | — | — | — | — |
| Gemini 2.0 Flash | 76.5 | — | 58.6 | — | — | — | — | — | — |
| Gemini 1.5 Pro | 73.1 | — | — | — | 57.5 | 84.1 | — | — | — |
| Gemini 1.5 Flash | — | — | — | — | — | — | — | — | — |
| Llama 3.1 405B Instruct | 73.3 | — | 51.1 | — | — | — | — | — | — |
| Llama 3.1 70B Instruct | 65.5 | — | — | — | 48.6 | 80.5 | — | — | — |
| Llama 3.1 8B Instruct | — | — | 28.6 | — | — | — | — | — | — |
| Llama 3.3 70B Instruct | 71.2 | — | 49.5 | — | — | — | — | — | — |
| DeepSeek R1 | 84 | — | 71.5 | 79.8 | — | — | — | — | — |
| DeepSeek V3 | 75.9 | — | 60.9 | — | — | — | 85.5 | — | — |
| DeepSeek V2.5 | 66.2 | — | — | — | — | — | — | — | — |
| Mistral Large 2 | 66.8 | — | — | — | — | — | — | — | — |
| Codestral 22B | — | — | — | — | — | 81.1 | — | — | — |
| Qwen 2.5 72B Instruct | 71.1 | — | 44.6 | — | 52.3 | — | 81.2 | — | — |
| Qwen 2.5 Coder 32B Instruct | — | — | — | — | — | 92.7 | — | 54.2 | — |
| Grok 2 | — | — | — | — | — | — | — | — | — |
| GPT-6 Astra | — | — | 94.4 | — | — | — | — | — | — |
| GPT-5.6 Sol | 89.1 | 96.2 | 91.9 | — | 81.1 | — | — | — | — |
| GPT-5.6 Terra | — | — | 89 | — | 77.9 | — | — | — | — |
| GPT-5.6 Luna | — | — | 87.7 | — | 73.6 | — | — | — | — |
| Claude Fable 5.1 | 92.38 | — | 90.9 | — | 83.4 | — | — | — | — |
| Claude Opus 5 | 91.59 | 97 | 86.8 | — | 80.1 | — | — | — | — |
| Claude Sonnet 5 | — | — | 82.2 | — | 76 | — | — | — | — |
| Gemini 3.8 Flash | 90.22 | — | 95.3 | — | — | — | — | — | — |
| Gemini 3.7 Flash | 90.12 | — | 94.3 | — | 78.8 | — | — | — | — |
| Grok 4.6 | — | — | 93.3 | — | 78 | — | — | — | — |
| DeepSeek V4 Pro | — | 96.4 | 88.9 | — | 77.4 | — | — | — | — |
| DeepSeek V4 Flash | — | — | — | — | — | — | — | — | — |
| Gemini 3.1 Pro Preview | — | — | 94.4 | — | — | — | — | — | — |
| GPT-5.5 | — | — | 93.8 | — | — | — | — | — | — |
| GPT-5.6 Sol Pro | — | — | 93.8 | — | — | — | — | — | — |
| Gemini 3.6 Flash | — | — | 92.8 | — | — | — | — | — | — |
| Gemini 3.5 Flash | — | — | 92.8 | — | — | — | — | — | — |
| Kimi K3 | — | — | 91.5 | — | — | — | — | — | — |
| MiniMax M3 | — | — | 90.5 | — | — | — | — | — | — |
| GPT-5.4 | — | — | 90.3 | — | — | — | — | — | — |
| GPT-5.6 Luna Pro | — | — | 90.3 | — | — | — | — | — | — |
| Claude Opus 4.8 | — | — | 89.4 | — | — | — | — | — | — |
| Nova Micro 1.0 | — | — | 89.1 | — | — | — | — | — | — |
| Claude Opus 4.7 | — | — | 88.9 | — | — | — | — | — | — |
| Gemini 3 Flash Preview | — | — | 88.3 | — | — | — | — | — | — |
| DeepSeek V4 Flash Vision Exp | — | — | 88.2 | — | — | — | — | — | — |
| GPT-5.2 | — | — | 87.9 | — | — | — | — | — | — |
| GLM 5.3 Flash | — | — | 86.7 | — | — | — | — | — | — |
| DeepSeek V4 Flash 0423 | — | — | 86.6 | — | — | — | — | — | — |
| Claude Opus 4.5 | — | — | 86.6 | — | — | — | — | — | — |
| Inkling Small | — | — | 86.5 | — | — | — | — | — | — |
| DeepSeek V4 Pro 0423 | — | — | 86.4 | — | — | — | — | — | — |
| GPT-5.1 | — | — | 86.2 | — | — | — | — | — | — |
| GPT-5 | — | — | 86 | — | — | — | — | — | — |
| Qwen3.5 397B A17B | — | — | 85.9 | — | — | — | — | — | — |
| GLM 5.2 | — | — | 85.8 | — | — | — | — | — | — |
| Qwen3.8 2.4T A95B | — | — | 85.7 | — | — | — | — | — | — |
| Claude Opus 4.6 | — | — | 85.6 | — | — | — | — | — | — |
| DeepSeek V4 Flash 0731 | — | — | 85.2 | — | — | — | — | — | — |
| MiniMax M2.1 | — | — | 85.1 | — | — | — | — | — | — |
| Kimi K2.5 | — | — | 84.9 | — | — | — | — | — | — |
| GLM 5.3 | — | — | 84.8 | — | — | — | — | — | — |
| MiniMax M2.5 | — | — | 84.2 | — | — | — | — | — | — |
| MiMo-V2.5-Pro | — | — | 84.2 | — | — | — | — | — | — |
| MiniMax M2.7 | — | — | 84.1 | — | — | — | — | — | — |
| GPT-5.4 Mini | — | — | 83.7 | — | — | — | — | — | — |
| Kimi K2 Thinking | — | — | 83.6 | — | — | — | — | — | — |
| Qwen3.5-122B-A10B | — | — | 83.6 | — | — | — | — | — | — |
| Gemini 3.5 Flash Lite | — | — | 83.5 | — | — | — | — | — | — |
| GLM 4.7 | — | — | 83.5 | — | — | — | — | — | — |
| Gemma 4 31B | — | — | 83.4 | — | — | — | — | — | — |
| Kimi K2.6 | — | — | 83.3 | — | — | — | — | — | — |
| Inkling | — | — | 83.2 | — | — | — | — | — | — |
| GLM 5.1 | — | — | 82.9 | — | — | — | — | — | — |
| Claude Fable 5 | — | — | 82.8 | — | — | — | — | — | — |
| Claude Sonnet 4.5 | — | — | 82.6 | — | — | — | — | — | — |
| Nemotron 3 Ultra | — | — | 82.2 | — | — | — | — | — | — |
| DeepSeek V3.2 | — | — | 82 | — | — | — | — | — | — |
| Qwen3.8 27B | — | — | 81.8 | — | — | — | — | — | — |
| Qwen3.5-35B-A3B | — | — | 81.7 | — | — | — | — | — | — |
| Kimi K2.7 Code | — | — | 80.7 | — | — | — | — | — | — |
| MiMo-V2-Flash | — | — | 80.7 | — | — | — | — | — | — |
| Qwen3.6 35B A3B | — | — | 80.6 | — | — | — | — | — | — |
| GPT-5.3 Chat | — | — | 80.5 | — | — | — | — | — | — |
| DeepSeek V3.1 Terminus | — | — | 80.4 | — | — | — | — | — | — |
| Muse Glimmer 30B | — | — | 80.2 | — | — | — | — | — | — |
| Gemini 3.1 Flash Lite | — | — | 80.1 | — | — | — | — | — | — |
| Qwen3 235B A22B Thinking 2507 | — | — | 80 | — | — | — | — | — | — |
| Gemini 2.5 Pro | — | — | 79.9 | — | — | — | — | — | — |
| GPT-5.2 Chat | — | — | 79.8 | — | — | — | — | — | — |
| GPT-5 Mini | — | — | 79.6 | — | — | — | — | — | — |
| Qwen3.6 27B | — | — | 79.6 | — | — | — | — | — | — |
| DeepSeek V3.1 | — | — | 79.3 | — | — | — | — | — | — |
| R1 0528 | — | — | 79.1 | — | — | — | — | — | — |
| GLM 4.6 | — | — | 78.8 | — | — | — | — | — | — |
| DeepSeek V3.2 Exp | — | — | 78.8 | — | — | — | — | — | — |
| GPT-5.4 Nano | — | — | 77.9 | — | — | — | — | — | — |
| Qwen3.5-9B | — | — | 77.9 | — | — | — | — | — | — |
| GLM 5 | — | — | 77.5 | — | — | — | — | — | — |
| Claude Sonnet 4 | — | — | 76.7 | — | — | — | — | — | — |
| Step 3.7 Flash | — | — | 76.6 | — | — | — | — | — | — |
| MiMo-V2.5 | — | — | 76.3 | — | — | — | — | — | — |
| Kimi K2 0905 | — | — | 75.9 | — | — | — | — | — | — |
| Mistral Small 4 | — | — | 75.8 | — | — | — | — | — | — |
| gpt-oss-120b | — | — | 75.1 | — | — | — | — | — | — |
| Auto Router (Beta) | — | — | 74.7 | — | — | — | — | — | — |
| Qwen3 235B A22B Instruct 2507 | — | — | 74.4 | — | — | — | — | — | — |
| Qwen3 Coder Next | — | — | 74.1 | — | — | — | — | — | — |
| Gemma 4 26B A4B | — | — | 73.9 | — | — | — | — | — | — |
| Ling 3.0 Flash | — | — | 73 | — | — | — | — | — | — |
| Gemini 2.5 Flash | — | — | 72.7 | — | — | — | — | — | — |
| Claude Haiku 4.5 | — | — | 72.1 | — | — | — | — | — | — |
| GPT-5 Nano | — | — | 70.9 | — | — | — | — | — | — |
| Qwen3 Next 80B A3B Instruct | — | — | 70.7 | — | — | — | — | — | — |
| Qwen3 VL 235B A22B Instruct | — | — | 70.5 | — | — | — | — | — | — |
| Nemotron 3.5 Lightning | — | — | 69 | — | — | — | — | — | — |
| Gemma 4 26B A4B IT (free) | — | — | 68.5 | — | — | — | — | — | — |
| gpt-oss-20b | — | — | 66.6 | — | — | — | — | — | — |
| Llama 4 Maverick | — | — | 65.9 | — | — | — | — | — | — |
| GPT-4.1 | — | — | 64.9 | — | — | — | — | — | — |
| Qwen3 VL 30B A3B Instruct | — | — | 64.8 | — | — | — | — | — | — |
| GLM 4.5 Air | — | — | 64.5 | — | — | — | — | — | — |
| GPT-4.1 Mini | — | — | 64.1 | — | — | — | — | — | — |
| Qwen3 30B A3B Instruct 2507 | — | — | 63.9 | — | — | — | — | — | — |
| Qwen3 Coder 480B A35B | — | — | 61.7 | — | — | — | — | — | — |
| Qwen3 30B A3B | — | — | 61.4 | — | — | — | — | — | — |
| Qwen3 32B | — | — | 61.4 | — | — | — | — | — | — |
| Nemotron 3 Nano 30B A3B | — | — | 60.7 | — | — | — | — | — | — |
| Qwen3 14B | — | — | 59.4 | — | — | — | — | — | — |
| DeepSeek V3 0324 | — | — | 58.1 | — | — | — | — | — | — |
| Gemini 2.5 Flash Lite | — | — | 55 | — | — | — | — | — | — |
| Claude Sonnet 4.6 | — | — | 54.4 | — | — | — | — | — | — |
| Qwen3 Coder 30B A3B Instruct | — | — | 52 | — | — | — | — | — | — |
| Qwen3 VL 8B Instruct | — | — | 51.8 | — | — | — | — | — | — |
| GPT-4o (2024-08-06) | — | — | 51.7 | — | — | — | — | — | — |
| GLM 4.7 Flash | — | — | 51.4 | — | — | — | — | — | — |
| GPT-4.1 Nano | — | — | 50.9 | — | — | — | — | — | — |
| GPT-4o (2024-05-13) | — | — | 50.7 | — | — | — | — | — | — |
| Qwen2.5 VL 72B Instruct | — | — | 44.4 | — | — | — | — | — | — |
| Qwen2.5 7B Instruct | — | — | 32.8 | — | — | — | — | — | — |
| Mistral Nemo | — | — | 31.6 | — | — | — | — | — | — |
| Llama 3 8B Lunaris | — | — | 26.9 | — | — | — | — | — | — |
| Ling 3.0 Flash Fin | — | — | 19.7 | — | — | — | — | — | — |
| Llama 3.2 3B Instruct | — | — | 11.6 | — | — | — | — | — | — |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | — | — | 10.3 | — | — | — | — | — | — |
Dash means no non-superseded observation. This is evidence coverage, not a quality score.
Benchmark versioning & contamination policy
AI Bench Index treats benchmark suites as evolving empirical measurements rather than timeless truths. When a benchmark suite updates its questions, fixes labeling errors, or releases a sanitized subset (e.g. SWE-bench Verified vs SWE-bench Lite), we preserve version tags on each observation. Older scores are never blended with newer test harness executions.
