Direct publisher snapshots
Published AI benchmark snapshots
Compare source-linked results while preserving each publisher’s model labels, configurations, dates and scoring rules. Cross-board ranks are not treated as interchangeable. Only scores stated in the publisher page’s rendered summary are reproduced; missing rows are never filled from model cards, estimates, or mixed settings.
Catalogue review recorded · individual claims have separate datesStructured datasetSourceMethodology
GPQA Diamond
Source: OpenRouterOpenRouter default-routing board
GPQA Diamond is a graduate-level multiple-choice benchmark in biology, physics, and chemistry. OpenRouter runs the same fixed question set across provider endpoints with default routing to compare capability, routing, and the practical cost of solving difficult scientific problems.
| Rank | Model | Provider | Score |
|---|---|---|---|
| #1 | Google DeepMind | 94.4% | |
| #2 | OpenAI | 94.4% | |
| #3 | Google DeepMind | 94.3% | |
| #4 | OpenAI | 93.8% | |
| #5 | OpenAI | 93.8% | |
| #6 | xAI | 93.3% | |
| #7 | Google DeepMind | 92.8% | |
| #8 | Google DeepMind | 92.8% | |
| #9 | OpenAI | 91.9% | |
| #10 | Moonshot AI | 91.5% | |
| #11 | Anthropic | 90.9% | |
| #12 | MiniMax | 90.5% | |
| #13 | OpenAI | 90.3% | |
| #14 | OpenAI | 90.3% | |
| #15 | Anthropic | 89.4% | |
| #16 | Amazon | 89.1% | |
| #17 | OpenAI | 89.0% | |
| #18 | Anthropic | 88.9% | |
| #19 | DeepSeek | 88.9% | |
| #20 | Google DeepMind | 88.3% | |
| #21 | DeepSeek | 88.2% | |
| #22 | OpenAI | 87.9% | |
| #23 | OpenAI | 87.7% | |
| #24 | Anthropic | 86.8% | |
| #25 | Z.ai | 86.7% | |
| #26 | DeepSeek | 86.6% | |
| #27 | Anthropic | 86.6% | |
| #28 | Thinking Machines | 86.5% | |
| #29 | DeepSeek | 86.4% | |
| #30 | OpenAI | 86.2% |
Show remaining 99 of 129 rows from the source board
| Rank | Model | Provider | Score |
|---|---|---|---|
| #31 | OpenAI | 86.0% | |
| #32 | Alibaba Cloud | 85.9% | |
| #33 | Z.ai | 85.8% | |
| #34 | Alibaba Cloud | 85.7% | |
| #35 | Anthropic | 85.6% | |
| #36 | DeepSeek | 85.2% | |
| #37 | MiniMax | 85.1% | |
| #38 | Moonshot AI | 84.9% | |
| #39 | Z.ai | 84.8% | |
| #40 | MiniMax | 84.2% | |
| #41 | Xiaomi | 84.2% | |
| #42 | MiniMax | 84.1% | |
| #43 | OpenAI | 83.7% | |
| #44 | Moonshot AI | 83.6% | |
| #45 | Alibaba Cloud | 83.6% | |
| #46 | Google DeepMind | 83.5% | |
| #47 | Z.ai | 83.5% | |
| #48 | Google DeepMind | 83.4% | |
| #49 | Moonshot AI | 83.3% | |
| #50 | Thinking Machines | 83.2% | |
| #51 | Z.ai | 82.9% | |
| #52 | Anthropic | 82.8% | |
| #53 | Anthropic | 82.6% | |
| #54 | Anthropic | 82.2% | |
| #55 | NVIDIA | 82.2% | |
| #56 | DeepSeek | 82.0% | |
| #57 | Alibaba Cloud | 81.8% | |
| #58 | Alibaba Cloud | 81.7% | |
| #59 | Moonshot AI | 80.7% | |
| #60 | Xiaomi | 80.7% | |
| #61 | Alibaba Cloud | 80.6% | |
| #62 | OpenAI | 80.5% | |
| #63 | DeepSeek | 80.4% | |
| #64 | Meta AI | 80.2% | |
| #65 | Google DeepMind | 80.1% | |
| #66 | Alibaba Cloud | 80.0% | |
| #67 | Google DeepMind | 79.9% | |
| #68 | OpenAI | 79.8% | |
| #69 | OpenAI | 79.6% | |
| #70 | Alibaba Cloud | 79.6% | |
| #71 | DeepSeek | 79.3% | |
| #72 | DeepSeek | 79.1% | |
| #73 | Z.ai | 78.8% | |
| #74 | DeepSeek | 78.8% | |
| #75 | OpenAI | 77.9% | |
| #76 | Alibaba Cloud | 77.9% | |
| #77 | Z.ai | 77.5% | |
| #78 | Anthropic | 76.7% | |
| #79 | StepFun | 76.6% | |
| #80 | Xiaomi | 76.3% | |
| #81 | Moonshot AI | 75.9% | |
| #82 | Mistral AI | 75.8% | |
| #83 | OpenAI | 75.1% | |
| #84 | Auto Router (Beta) | OpenRouter | 74.7% |
| #85 | Alibaba Cloud | 74.4% | |
| #86 | Alibaba Cloud | 74.1% | |
| #87 | Google DeepMind | 73.9% | |
| #88 | inclusionAI | 73.0% | |
| #89 | Google DeepMind | 72.7% | |
| #90 | Anthropic | 72.1% | |
| #91 | OpenAI | 70.9% | |
| #92 | Alibaba Cloud | 70.7% | |
| #93 | Alibaba Cloud | 70.5% | |
| #94 | NVIDIA | 69.0% | |
| #95 | Google DeepMind | 68.5% | |
| #96 | OpenAI | 66.6% | |
| #97 | Meta AI | 65.9% | |
| #98 | OpenAI | 64.9% | |
| #99 | Alibaba Cloud | 64.8% | |
| #100 | Z.ai | 64.5% | |
| #101 | OpenAI | 64.1% | |
| #102 | Alibaba Cloud | 63.9% | |
| #103 | Alibaba Cloud | 61.7% | |
| #104 | Alibaba Cloud | 61.4% | |
| #105 | Alibaba Cloud | 61.4% | |
| #106 | DeepSeek | 60.9% | |
| #107 | NVIDIA | 60.7% | |
| #108 | Alibaba Cloud | 59.4% | |
| #109 | DeepSeek | 58.1% | |
| #110 | Google DeepMind | 55.0% | |
| #111 | Anthropic | 54.4% | |
| #112 | Alibaba Cloud | 52.0% | |
| #113 | Alibaba Cloud | 51.8% | |
| #114 | OpenAI | 51.7% | |
| #115 | Z.ai | 51.4% | |
| #116 | OpenAI | 50.9% | |
| #117 | OpenAI | 50.7% | |
| #118 | OpenAI | 50.3% | |
| #119 | Meta AI | 49.5% | |
| #120 | Alibaba Cloud | 44.6% | |
| #121 | Alibaba Cloud | 44.4% | |
| #122 | OpenAI | 43.4% | |
| #123 | Alibaba Cloud | 32.8% | |
| #124 | Mistral AI | 31.6% | |
| #125 | Meta AI | 28.6% | |
| #126 | Sao10K | 26.9% | |
| #127 | inclusionAI | 19.7% | |
| #128 | Meta AI | 11.6% | |
| #129 | Google DeepMind | 10.3% |
Source summary: “Gemini 3.1 Pro Preview ranks #1 and GPT-6 Astra #2, both at 94.4%, with Gemini 3.7 Flash third at 94.3% on OpenRouter's default-routing board.”
Publisher page renders a 129-row default-routing leaderboard; all 129 source-rendered rows are copied here in the publisher's own order. Provider-pinned sub-rows are not copied. Checked .
Artificial Analysis Intelligence Index (v4.2 cohort)
Source: Artificial Analysisv4.2
A composite benchmark aggregating ten challenging evaluations to provide a holistic measure of AI capabilities across mathematics, science, coding, and reasoning.
| Rank | Model | Provider | Score |
|---|---|---|---|
| #1 | Anthropic | 57 | |
| #2 | OpenAI | 55 | |
| #3 | OpenAI | 54 |
Source summary: “Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) scores the highest on Artificial Analysis Intelligence Index with a score of 57, followed by GPT-6 Astra (max) with a score of 55, and GPT-6 Astra (xhigh) with a score of 54”
Publisher page reports 27 displayed results of 630 models and includes an ‘Estimate (independent evaluation forthcoming)’ label; only its source-rendered top-three summary is copied here. Historical v4.2 cohort: superseded by the publisher's v4.3 edition (2026-09-07) and kept unmodified for edition history. Checked .
Artificial Analysis Intelligence Index (v4.3 cohort)
Source: Artificial Analysisv4.3
Current-edition capture of the composite index (aggregating ten evaluations across agents, coding, general capability and scientific reasoning) under the publisher's v4.3 methodology, which changed agentic evaluation components. Scores are NOT comparable with the v4.2 cohort above: a methodology change is not a model change.
| Rank | Model | Provider | Score |
|---|---|---|---|
| #1 | Anthropic | 53 | |
| #2 | Anthropic | 53 | |
| #3 | OpenAI | 53 |
Source summary: “Under v4.3, the publisher's rendered top three are tied at index 53: Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback), Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) and GPT-6 Astra (max), in the publisher's rendered order.”
Publisher page reports 28 displayed results of 633 models; only its source-rendered top-three summary is copied here. The rendered order lists the tied rows 1–3; the tie is the publisher's rendering, not our resolution. Checked .
Terminal-Bench v2.1
Source: Artificial AnalysisTerminus 2 · pass@1 · 3 repeats
A verified refresh of Terminal-Bench v2.0 — 89 curated tasks across software engineering, system administration, data processing, model training, and security, with environment and instruction fixes so scores reflect agent capability rather than environment gaps.
| Rank | Model | Provider | Score |
|---|---|---|---|
| #1 | Anthropic | 91.4% | |
| #2 | Anthropic | 91.0% | |
| #3 | OpenAI | 89.9% |
Source summary: “Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) scores the highest on Terminal-Bench v2.1 with a score of 91.4%, followed by Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) with a score of 91.0%, and GPT-6 Astra (high) with a score of 89.9%”
Publisher page reports 27 displayed results of 235 models; only its source-rendered top-three summary is copied here. Checked .
MMLU-Pro
Source: Artificial AnalysisArtificial Analysis current board
An enhanced version of MMLU with 12,000 graduate-level questions across 14 subject areas, featuring ten answer options and deeper reasoning requirements.
| Rank | Model | Provider | Score |
|---|---|---|---|
| #1 | 89.8% | ||
| =2 | 89.5% | ||
| =2 | Anthropic | 89.5% |
Source summary: “Gemini 3 Pro Preview (high) scores the highest on MMLU-Pro with a score of 89.8%, followed by Gemini 3 Pro Preview (low) with a score of 89.5%, and Claude Opus 4.5 (Reasoning) with a score of 89.5%”
Publisher page reports 6 displayed results of 349 models; only its source-rendered top-three summary is copied here. Checked .
Boards linked without copied scores
These sources did not expose ranking rows in server-rendered content, so this site does not guess or preserve an unchecked local copy.
- LiveBench
The public page requires client-side JavaScript and did not expose a source-rendered ranking table to this snapshot process.
- SWE-bench Verified
The source-rendered page documents the benchmark and harness variants but does not expose its interactive ranking rows.
