Benchmark evidence

AI benchmark leaderboards

Browse benchmark results by the organisations and research labs that publish them: Cursor, LMSYS Arena, Artificial Analysis, DeepSWE, SWE-bench Verified, LiveBench, and Provider Catalogs. Compare Pareto efficiency (accuracy vs cost/tokens/steps) and raw score tables.

Catalogue review recorded · individual claims have separate datesStructured datasetSourceMethodology

Swipe sideways to see all publishers.

Highest displayed publisher result

Claude Fable 5.1

66 (Index score)

Rows shown

25 rows

Index score

Snapshot date

2026-09-08

Cohort edition v4.1.1 era — historical rows stay under their captured edition
Artificial Analysis Intelligence Index

Evaluation Board:

Understanding Intelligence Index

A composite index aggregating several evaluations into a comparative model-quality signal.

What it reveals

It is useful as one broad reference point when a buyer needs an initial shortlist.

What it cannot establish

Composite scores hide task-specific variance, version changes, and the trade-off between quality and cost.

These are transcribed publisher results, not tests run by AI Bench Index. A publisher’s order may include different model settings. Row sources and notes appear in the full table; the snapshot date is not a test-run date.

Read the publisher’s methodology and results ↗
Notice: Composite index for reasoning, coding, STEM knowledge, and autonomous agent tasks from Artificial Analysis.
Notice: Legacy pre-v4.2 cohort: values reconciled against the live board on 2026-08-29 (the v4.1.1 era; v4.2 launched 2026-09-04 and v4.3 2026-09-07 per the publisher changelog). Rows were re-provenanced on 2026-09-08 against dated Internet Archive captures of the same era; see the data-file header. The cohort is kept as captured and is not relabelled; v4.2 and v4.3 cohorts are captured separately on the leaderboards page.

Score profile

Index score by published row

Higher is better

#1Claude Fable 5.1
66
#2Claude Fable 5.1 (xhigh)
65
#3Claude Opus 5 (max)
63
#3Claude Opus 5 (xhigh)
63
#5Claude Fable 5 (with fallback)
62
#6Claude Opus 5 (high)
61
#6GPT-5.6 Sol (max)
61
#6GPT-6 Astra (max)
61
#6Grok 4.6 (high)
61
#10GLM-5.3 (max)
60
#10Grok 4.6 (xhigh)
60
#10Kimi K3 (max)
60
#13Claude Opus 5 (medium)
59
#13Gemini 3.8 Flash (high)
59
#13GPT-5.6 Sol (xhigh)
59
#13Grok 4.6 (medium)
59
#17Qwen3.8 2.4T A95B
58
#17Qwen3.8 Max
58
#19Gemini 3.8 Flash (medium)
57
#19GLM-5.3-Flash
57
#19GPT-5.6 Sol (high)
57
#19GPT-5.6 Terra (max)
57
#19Muse Spark 1.2 (xhigh)
57
#24Gemini 3.7 Flash (high)
56
#24GPT-5.6 Sol (medium)
56

Artificial Analysis composite intelligence index across leading frontier models.

Showing 25 of 614 transcribed rows

Catalogued suites

Standardized evaluation suites

Dataset definitions, contamination concerns, and recorded evaluation methods. Historical suites remain available as reference.

Reasoningv1.0

MMLU-Pro

Owner: TIGER AI Lab (Waterloo)academic

Massive Multitask Language Understanding Professional: 12,000+ complex reasoning questions across 14 disciplines with 10 answer options, which reduce (but do not eliminate) the chance of a correct random guess.

Measures
Hard multiple-choice questions across 14 academic subjects with 10 options.
Does not test
Tool use, long-horizon coding, or real repository patches.
Misread
Not interchangeable with original 5-option MMLU.

No pooled distribution: recorded evaluation conditions differ or are missing.

Metric: percentage
Direction: Higher is better
Observations: 18
Checked: Checked 4 Sep 2026
Comparability notes

Dataset 2024-06-01. Contamination low. Scores are never blended across harness versions.

Source: arXiv:2406.01574Methodology
Codingv2024.08

SWE-bench Verified

Owner: Princeton NLP & OpenAIindustry

Human-validated subset of 500 real GitHub issue resolution tasks evaluating end-to-end software engineering on production codebases.

Measures
Resolving real GitHub issues on a 500-task human-validated subset.
Does not test
Chat coding quality in an IDE, or SWE-bench Pro / multilingual variants.
Misread
Vals mini-swe-agent 2026 rows are not the same experiment as 2024 two-tool or Aider scaffolds.
Metric: pass@1
Direction: Higher is better
Observations: 11
Checked: Checked 4 Sep 2026
Comparability notes

Dataset 2024-08-15. Contamination high. Scores are never blended across harness versions.

Source: arXiv:2310.06770Methodology
Reasoningv1.0

GPQA Diamond

Owner: NYU / Anthropic / ARCacademic

Google-Proof Q&A: 198 graduate-level biology, chemistry, and physics questions authored and validated by PhD domain experts.

Measures
198 graduate-level biology, chemistry, and physics questions written by domain experts.
Does not test
Laboratory skill, tool use, or coding.
Misread
Diamond is a subset of GPQA; do not mix Diamond with the full set.

No pooled distribution: recorded evaluation conditions differ or are missing.

Metric: percentage
Direction: Higher is better
Observations: 147
Checked: Checked 7 Sep 2026
Comparability notes

Dataset 2023-11-20. Contamination low. Scores are never blended across harness versions.

Source: arXiv:2311.12022Methodology
Mathematicsv1.0

MATH 500

superseded suite · retained as historical reference, not an active frontier ranking.

Owner: OpenAI / Hendrycksacademic

Representative 500-problem subset of challenging competition-level high school mathematics from the Hendrycks MATH benchmark.

Measures
A 500-problem slice of Hendrycks MATH competition items.
Does not test
Proof writing or MATH-500 replacements such as AIME live contests.
Misread
This suite is catalogued as superseded for active leaderboards.

No pooled distribution: recorded evaluation conditions differ or are missing.

Metric: percentage
Direction: Higher is better
Observations: 7
Checked: Checked 1 Sep 2026
Comparability notes

Dataset 2024-05-01. Contamination moderate. Scores are never blended across harness versions.

Source: arXiv:2103.03874Methodology
Mathematicsv2024

AIME 2024

Owner: Mathematical Association of Americaindustry

American Invitational Mathematics Examination 2024: 30 non-multiple choice competition questions for the top 5% of math students.

Measures
AIME 2024 contest problems (integer answers).
Does not test
Open-ended proof contests or MATH-500.
Misread
AIME 2024 is not AIME 2025.

No pooled distribution: recorded evaluation conditions differ or are missing.

Metric: percentage
Direction: Higher is better
Observations: 4
Checked: Checked 1 Sep 2026
Comparability notes

Dataset 2024-02-15. Contamination low. Scores are never blended across harness versions.

Source: MAA AIMEMethodology
Reasoningv2024.10

LiveBench

Owner: Abacus.ai & NYUacademic

Contamination-resistant benchmark with questions sourced from post-cutoff competitions, arXiv papers, and real news. LiveBench was introduced with monthly question releases; the current public site describes a six-month refresh cadence. Results are only compared within the same release and configuration.

Measures
A frequently refreshed mix of tasks intended to reduce contamination.
Does not test
A single skill such as only coding or only math.
Misread
Snapshots from different dates are not one ladder.

No pooled distribution: recorded evaluation conditions differ or are missing.

Metric: percentage
Direction: Higher is better
Observations: 15
Checked: Checked 4 Sep 2026
Comparability notes

Dataset 2024-10-01. Contamination low. Scores are never blended across harness versions.

Source: arXiv:2406.19314Methodology
Codingv1.0

HumanEval

Owner: OpenAIacademic

164 hand-crafted Python programming problems with unit tests measuring functional correctness (pass@1) from docstrings.

Measures
Short Python function synthesis from docstrings (pass@k).
Does not test
Multi-file repository repair (that is SWE-bench).
Misread
HumanEval is not SWE-bench Verified.

No pooled distribution: recorded evaluation conditions differ or are missing.

Metric: pass@1
Direction: Higher is better
Observations: 6
Checked: Checked 1 Sep 2026
Comparability notes

Dataset 2021-07-07. Contamination high. Scores are never blended across harness versions.

Source: arXiv:2107.03374Methodology
Reasoningv0.1

Arena-Hard Auto

Owner: LMSYS Orgacademic

500 challenging user queries from Chatbot Arena evaluated with automated LLM judges, showing 89.1% correlation with crowdsourced human votes.

Measures
Hard conversational prompts judged against a baseline policy.
Does not test
Deterministic unit-tested coding.
Misread
LMSYS Arena Elo boards are separate from Arena-Hard.

Median 85.2 across 5 models with evidence (range 79.3–86.8).

Metric: percentage
Direction: Higher is better
Observations: 5
Checked: Checked 1 Sep 2026
Comparability notes

Dataset 2024-06-15. Contamination low. Scores are never blended across harness versions.

Source: arXiv:2406.11939Methodology
Codingv1.0

BigCodeBench

Owner: BigCode Projectacademic

1,140 challenging Python programming tasks invoking 139 external libraries, evaluating real-world software development.

Measures
More realistic Python programming tasks than HumanEval.
Does not test
GitHub issue resolution on existing repos.
Misread
Instruct vs complete splits are not interchangeable.
Metric: pass@1
Direction: Higher is better
Observations: 1
Checked: Checked 1 Sep 2026
Comparability notes

Dataset 2024-06-25. Contamination low. Scores are never blended across harness versions.

Source: arXiv:2406.15877Methodology
Codingv1.0

SWE-bench Lite

Owner: Princeton NLPindustry

300 issue resolution tasks selected for clean self-contained evaluation across 12 major Python repositories.

Measures
A smaller SWE-bench subset than Verified.
Does not test
The 500-task Verified set.
Misread
Lite is not Verified.
Metric: pass@1
Direction: Higher is better
Observations: 2
Checked: Checked 1 Sep 2026
Comparability notes

Dataset 2024-03-01. Contamination low. Scores are never blended across harness versions.

Source: arXiv:2310.06770Methodology

Coverage matrix

Which tracked models have a non-superseded receipt on each active suite. Missing cells are missing evidence, not zeros.

ModelMMLU-ProSWE-bench VerifiedGPQA DiamondAIME 2024LiveBenchHumanEvalArena-Hard AutoBigCodeBenchSWE-bench Lite
OpenAI o175.779.268.286.8
OpenAI o1-mini
GPT-4o72.650.39.359.290.279.338
GPT-4o mini62.543.4
Claude 3.5 Sonnet78651664.893.785.243.3
Claude 3.5 Haiku
Claude 3 Opus
Gemini 2.0 Flash76.558.6
Gemini 1.5 Pro73.157.584.1
Gemini 1.5 Flash
Llama 3.1 405B Instruct73.351.1
Llama 3.1 70B Instruct65.548.680.5
Llama 3.1 8B Instruct28.6
Llama 3.3 70B Instruct71.249.5
DeepSeek R18471.579.8
DeepSeek V375.960.985.5
DeepSeek V2.566.2
Mistral Large 266.8
Codestral 22B81.1
Qwen 2.5 72B Instruct71.144.652.381.2
Qwen 2.5 Coder 32B Instruct92.754.2
Grok 2
GPT-6 Astra94.4
GPT-5.6 Sol89.196.291.981.1
GPT-5.6 Terra8977.9
GPT-5.6 Luna87.773.6
Claude Fable 5.192.3890.983.4
Claude Opus 591.599786.880.1
Claude Sonnet 582.276
Gemini 3.8 Flash90.2295.3
Gemini 3.7 Flash90.1294.378.8
Grok 4.693.378
DeepSeek V4 Pro96.488.977.4
DeepSeek V4 Flash
Gemini 3.1 Pro Preview94.4
GPT-5.593.8
GPT-5.6 Sol Pro93.8
Gemini 3.6 Flash92.8
Gemini 3.5 Flash92.8
Kimi K391.5
MiniMax M390.5
GPT-5.490.3
GPT-5.6 Luna Pro90.3
Claude Opus 4.889.4
Nova Micro 1.089.1
Claude Opus 4.788.9
Gemini 3 Flash Preview88.3
DeepSeek V4 Flash Vision Exp88.2
GPT-5.287.9
GLM 5.3 Flash86.7
DeepSeek V4 Flash 042386.6
Claude Opus 4.586.6
Inkling Small86.5
DeepSeek V4 Pro 042386.4
GPT-5.186.2
GPT-586
Qwen3.5 397B A17B85.9
GLM 5.285.8
Qwen3.8 2.4T A95B85.7
Claude Opus 4.685.6
DeepSeek V4 Flash 073185.2
MiniMax M2.185.1
Kimi K2.584.9
GLM 5.384.8
MiniMax M2.584.2
MiMo-V2.5-Pro84.2
MiniMax M2.784.1
GPT-5.4 Mini83.7
Kimi K2 Thinking83.6
Qwen3.5-122B-A10B83.6
Gemini 3.5 Flash Lite83.5
GLM 4.783.5
Gemma 4 31B83.4
Kimi K2.683.3
Inkling83.2
GLM 5.182.9
Claude Fable 582.8
Claude Sonnet 4.582.6
Nemotron 3 Ultra82.2
DeepSeek V3.282
Qwen3.8 27B81.8
Qwen3.5-35B-A3B81.7
Kimi K2.7 Code80.7
MiMo-V2-Flash80.7
Qwen3.6 35B A3B80.6
GPT-5.3 Chat80.5
DeepSeek V3.1 Terminus80.4
Muse Glimmer 30B80.2
Gemini 3.1 Flash Lite80.1
Qwen3 235B A22B Thinking 250780
Gemini 2.5 Pro79.9
GPT-5.2 Chat79.8
GPT-5 Mini79.6
Qwen3.6 27B79.6
DeepSeek V3.179.3
R1 052879.1
GLM 4.678.8
DeepSeek V3.2 Exp78.8
GPT-5.4 Nano77.9
Qwen3.5-9B77.9
GLM 577.5
Claude Sonnet 476.7
Step 3.7 Flash76.6
MiMo-V2.576.3
Kimi K2 090575.9
Mistral Small 475.8
gpt-oss-120b75.1
Auto Router (Beta)74.7
Qwen3 235B A22B Instruct 250774.4
Qwen3 Coder Next74.1
Gemma 4 26B A4B73.9
Ling 3.0 Flash73
Gemini 2.5 Flash72.7
Claude Haiku 4.572.1
GPT-5 Nano70.9
Qwen3 Next 80B A3B Instruct70.7
Qwen3 VL 235B A22B Instruct70.5
Nemotron 3.5 Lightning69
Gemma 4 26B A4B IT (free)68.5
gpt-oss-20b66.6
Llama 4 Maverick65.9
GPT-4.164.9
Qwen3 VL 30B A3B Instruct64.8
GLM 4.5 Air64.5
GPT-4.1 Mini64.1
Qwen3 30B A3B Instruct 250763.9
Qwen3 Coder 480B A35B61.7
Qwen3 30B A3B61.4
Qwen3 32B61.4
Nemotron 3 Nano 30B A3B60.7
Qwen3 14B59.4
DeepSeek V3 032458.1
Gemini 2.5 Flash Lite55
Claude Sonnet 4.654.4
Qwen3 Coder 30B A3B Instruct52
Qwen3 VL 8B Instruct51.8
GPT-4o (2024-08-06)51.7
GLM 4.7 Flash51.4
GPT-4.1 Nano50.9
GPT-4o (2024-05-13)50.7
Qwen2.5 VL 72B Instruct44.4
Qwen2.5 7B Instruct32.8
Mistral Nemo31.6
Llama 3 8B Lunaris26.9
Ling 3.0 Flash Fin19.7
Llama 3.2 3B Instruct11.6
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)10.3

Dash means no non-superseded observation. This is evidence coverage, not a quality score.

Benchmark versioning & contamination policy

AI Bench Index treats benchmark suites as evolving empirical measurements rather than timeless truths. When a benchmark suite updates its questions, fixes labeling errors, or releases a sanitized subset (e.g. SWE-bench Verified vs SWE-bench Lite), we preserve version tags on each observation. Older scores are never blended with newer test harness executions.