Direct publisher snapshots

Published AI benchmark snapshots

Compare source-linked results while preserving each publisher’s model labels, configurations, dates and scoring rules. Cross-board ranks are not treated as interchangeable. Only scores stated in the publisher page’s rendered summary are reproduced; missing rows are never filled from model cards, estimates, or mixed settings.

Catalogue review recorded · individual claims have separate datesStructured datasetSourceMethodology

Publication rule: a score is shown only when the named source page states the model, rank, and score together. These are publisher snapshots, not AIBenchIndex measurements. Open the linked source for the full current interactive board.

GPQA Diamond

Source: OpenRouter

OpenRouter default-routing board

Publisher source

GPQA Diamond is a graduate-level multiple-choice benchmark in biology, physics, and chemistry. OpenRouter runs the same fixed question set across provider endpoints with default routing to compare capability, routing, and the practical cost of solving difficult scientific problems.

Publisher-rendered ranking rows
RankModelProviderScore
#1Gemini 3.1 Pro PreviewGoogle DeepMind94.4%
#2GPT-6 AstraOpenAI94.4%
#3Gemini 3.7 FlashGoogle DeepMind94.3%
#4GPT-5.5OpenAI93.8%
#5GPT-5.6 Sol ProOpenAI93.8%
#6Grok 4.6xAI93.3%
#7Gemini 3.6 FlashGoogle DeepMind92.8%
#8Gemini 3.5 FlashGoogle DeepMind92.8%
#9GPT-5.6 SolOpenAI91.9%
#10Kimi K3Moonshot AI91.5%
#11Claude Fable 5.1Anthropic90.9%
#12MiniMax M3MiniMax90.5%
#13GPT-5.4OpenAI90.3%
#14GPT-5.6 Luna ProOpenAI90.3%
#15Claude Opus 4.8Anthropic89.4%
#16Nova Micro 1.0Amazon89.1%
#17GPT-5.6 TerraOpenAI89.0%
#18Claude Opus 4.7Anthropic88.9%
#19DeepSeek V4 ProDeepSeek88.9%
#20Gemini 3 Flash PreviewGoogle DeepMind88.3%
#21DeepSeek V4 Flash Vision ExpDeepSeek88.2%
#22GPT-5.2OpenAI87.9%
#23GPT-5.6 LunaOpenAI87.7%
#24Claude Opus 5Anthropic86.8%
#25GLM 5.3 FlashZ.ai86.7%
#26DeepSeek V4 Flash 0423DeepSeek86.6%
#27Claude Opus 4.5Anthropic86.6%
#28Inkling SmallThinking Machines86.5%
#29DeepSeek V4 Pro 0423DeepSeek86.4%
#30GPT-5.1OpenAI86.2%
Show remaining 99 of 129 rows from the source board
Publisher-rendered ranking rows
RankModelProviderScore
#31GPT-5OpenAI86.0%
#32Qwen3.5 397B A17BAlibaba Cloud85.9%
#33GLM 5.2Z.ai85.8%
#34Qwen3.8 2.4T A95BAlibaba Cloud85.7%
#35Claude Opus 4.6Anthropic85.6%
#36DeepSeek V4 Flash 0731DeepSeek85.2%
#37MiniMax M2.1MiniMax85.1%
#38Kimi K2.5Moonshot AI84.9%
#39GLM 5.3Z.ai84.8%
#40MiniMax M2.5MiniMax84.2%
#41MiMo-V2.5-ProXiaomi84.2%
#42MiniMax M2.7MiniMax84.1%
#43GPT-5.4 MiniOpenAI83.7%
#44Kimi K2 ThinkingMoonshot AI83.6%
#45Qwen3.5-122B-A10BAlibaba Cloud83.6%
#46Gemini 3.5 Flash LiteGoogle DeepMind83.5%
#47GLM 4.7Z.ai83.5%
#48Gemma 4 31BGoogle DeepMind83.4%
#49Kimi K2.6Moonshot AI83.3%
#50InklingThinking Machines83.2%
#51GLM 5.1Z.ai82.9%
#52Claude Fable 5Anthropic82.8%
#53Claude Sonnet 4.5Anthropic82.6%
#54Claude Sonnet 5Anthropic82.2%
#55Nemotron 3 UltraNVIDIA82.2%
#56DeepSeek V3.2DeepSeek82.0%
#57Qwen3.8 27BAlibaba Cloud81.8%
#58Qwen3.5-35B-A3BAlibaba Cloud81.7%
#59Kimi K2.7 CodeMoonshot AI80.7%
#60MiMo-V2-FlashXiaomi80.7%
#61Qwen3.6 35B A3BAlibaba Cloud80.6%
#62GPT-5.3 ChatOpenAI80.5%
#63DeepSeek V3.1 TerminusDeepSeek80.4%
#64Muse Glimmer 30BMeta AI80.2%
#65Gemini 3.1 Flash LiteGoogle DeepMind80.1%
#66Qwen3 235B A22B Thinking 2507Alibaba Cloud80.0%
#67Gemini 2.5 ProGoogle DeepMind79.9%
#68GPT-5.2 ChatOpenAI79.8%
#69GPT-5 MiniOpenAI79.6%
#70Qwen3.6 27BAlibaba Cloud79.6%
#71DeepSeek V3.1DeepSeek79.3%
#72R1 0528DeepSeek79.1%
#73GLM 4.6Z.ai78.8%
#74DeepSeek V3.2 ExpDeepSeek78.8%
#75GPT-5.4 NanoOpenAI77.9%
#76Qwen3.5-9BAlibaba Cloud77.9%
#77GLM 5Z.ai77.5%
#78Claude Sonnet 4Anthropic76.7%
#79Step 3.7 FlashStepFun76.6%
#80MiMo-V2.5Xiaomi76.3%
#81Kimi K2 0905Moonshot AI75.9%
#82Mistral Small 4Mistral AI75.8%
#83gpt-oss-120bOpenAI75.1%
#84Auto Router (Beta)OpenRouter74.7%
#85Qwen3 235B A22B Instruct 2507Alibaba Cloud74.4%
#86Qwen3 Coder NextAlibaba Cloud74.1%
#87Gemma 4 26B A4BGoogle DeepMind73.9%
#88Ling 3.0 FlashinclusionAI73.0%
#89Gemini 2.5 FlashGoogle DeepMind72.7%
#90Claude Haiku 4.5Anthropic72.1%
#91GPT-5 NanoOpenAI70.9%
#92Qwen3 Next 80B A3B InstructAlibaba Cloud70.7%
#93Qwen3 VL 235B A22B InstructAlibaba Cloud70.5%
#94Nemotron 3.5 LightningNVIDIA69.0%
#95Gemma 4 26B A4B IT (free)Google DeepMind68.5%
#96gpt-oss-20bOpenAI66.6%
#97Llama 4 MaverickMeta AI65.9%
#98GPT-4.1OpenAI64.9%
#99Qwen3 VL 30B A3B InstructAlibaba Cloud64.8%
#100GLM 4.5 AirZ.ai64.5%
#101GPT-4.1 MiniOpenAI64.1%
#102Qwen3 30B A3B Instruct 2507Alibaba Cloud63.9%
#103Qwen3 Coder 480B A35BAlibaba Cloud61.7%
#104Qwen3 30B A3BAlibaba Cloud61.4%
#105Qwen3 32BAlibaba Cloud61.4%
#106DeepSeek V3DeepSeek60.9%
#107Nemotron 3 Nano 30B A3BNVIDIA60.7%
#108Qwen3 14BAlibaba Cloud59.4%
#109DeepSeek V3 0324DeepSeek58.1%
#110Gemini 2.5 Flash LiteGoogle DeepMind55.0%
#111Claude Sonnet 4.6Anthropic54.4%
#112Qwen3 Coder 30B A3B InstructAlibaba Cloud52.0%
#113Qwen3 VL 8B InstructAlibaba Cloud51.8%
#114GPT-4o (2024-08-06)OpenAI51.7%
#115GLM 4.7 FlashZ.ai51.4%
#116GPT-4.1 NanoOpenAI50.9%
#117GPT-4o (2024-05-13)OpenAI50.7%
#118GPT-4oOpenAI50.3%
#119Llama 3.3 70B InstructMeta AI49.5%
#120Qwen 2.5 72B InstructAlibaba Cloud44.6%
#121Qwen2.5 VL 72B InstructAlibaba Cloud44.4%
#122GPT-4o miniOpenAI43.4%
#123Qwen2.5 7B InstructAlibaba Cloud32.8%
#124Mistral NemoMistral AI31.6%
#125Llama 3.1 8B InstructMeta AI28.6%
#126Llama 3 8B LunarisSao10K26.9%
#127Ling 3.0 Flash FininclusionAI19.7%
#128Llama 3.2 3B InstructMeta AI11.6%
#129Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)Google DeepMind10.3%

Source summary: “Gemini 3.1 Pro Preview ranks #1 and GPT-6 Astra #2, both at 94.4%, with Gemini 3.7 Flash third at 94.3% on OpenRouter's default-routing board.

Publisher page renders a 129-row default-routing leaderboard; all 129 source-rendered rows are copied here in the publisher's own order. Provider-pinned sub-rows are not copied. Checked .

Artificial Analysis Intelligence Index (v4.2 cohort)

Source: Artificial Analysis

v4.2

Publisher source

A composite benchmark aggregating ten challenging evaluations to provide a holistic measure of AI capabilities across mathematics, science, coding, and reasoning.

Publisher-rendered ranking rows
RankModelProviderScore
#1Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic57
#2GPT-6 Astra (max)OpenAI55
#3GPT-6 Astra (xhigh)OpenAI54

Source summary: “Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) scores the highest on Artificial Analysis Intelligence Index with a score of 57, followed by GPT-6 Astra (max) with a score of 55, and GPT-6 Astra (xhigh) with a score of 54

Publisher page reports 27 displayed results of 630 models and includes an ‘Estimate (independent evaluation forthcoming)’ label; only its source-rendered top-three summary is copied here. Historical v4.2 cohort: superseded by the publisher's v4.3 edition (2026-09-07) and kept unmodified for edition history. Checked .

Artificial Analysis Intelligence Index (v4.3 cohort)

Source: Artificial Analysis

v4.3

Publisher source

Current-edition capture of the composite index (aggregating ten evaluations across agents, coding, general capability and scientific reasoning) under the publisher's v4.3 methodology, which changed agentic evaluation components. Scores are NOT comparable with the v4.2 cohort above: a methodology change is not a model change.

Publisher-rendered ranking rows
RankModelProviderScore
#1Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic53
#2Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Anthropic53
#3GPT-6 Astra (max)OpenAI53

Source summary: “Under v4.3, the publisher's rendered top three are tied at index 53: Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback), Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) and GPT-6 Astra (max), in the publisher's rendered order.

Publisher page reports 28 displayed results of 633 models; only its source-rendered top-three summary is copied here. The rendered order lists the tied rows 1–3; the tie is the publisher's rendering, not our resolution. Checked .

Terminal-Bench v2.1

Source: Artificial Analysis

Terminus 2 · pass@1 · 3 repeats

Publisher source

A verified refresh of Terminal-Bench v2.0 — 89 curated tasks across software engineering, system administration, data processing, model training, and security, with environment and instruction fixes so scores reflect agent capability rather than environment gaps.

Publisher-rendered ranking rows
RankModelProviderScore
#1Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic91.4%
#2Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Anthropic91.0%
#3GPT-6 Astra (high)OpenAI89.9%

Source summary: “Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) scores the highest on Terminal-Bench v2.1 with a score of 91.4%, followed by Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) with a score of 91.0%, and GPT-6 Astra (high) with a score of 89.9%

Publisher page reports 27 displayed results of 235 models; only its source-rendered top-three summary is copied here. Checked .

MMLU-Pro

Source: Artificial Analysis

Artificial Analysis current board

Publisher source

An enhanced version of MMLU with 12,000 graduate-level questions across 14 subject areas, featuring ten answer options and deeper reasoning requirements.

Publisher-rendered ranking rows
RankModelProviderScore
#1Gemini 3 Pro Preview (high)Google89.8%
=2Gemini 3 Pro Preview (low)Google89.5%
=2Claude Opus 4.5 (Reasoning)Anthropic89.5%

Source summary: “Gemini 3 Pro Preview (high) scores the highest on MMLU-Pro with a score of 89.8%, followed by Gemini 3 Pro Preview (low) with a score of 89.5%, and Claude Opus 4.5 (Reasoning) with a score of 89.5%

Publisher page reports 6 displayed results of 349 models; only its source-rendered top-three summary is copied here. Checked .

Boards linked without copied scores

These sources did not expose ranking rows in server-rendered content, so this site does not guess or preserve an unchecked local copy.

  • LiveBench

    The public page requires client-side JavaScript and did not expose a source-rendered ranking table to this snapshot process.

  • SWE-bench Verified

    The source-rendered page documents the benchmark and harness variants but does not expose its interactive ranking rows.

Read benchmark methodology and limitations