Platform evaluation standards
Methodology & Transparency
Last updated: 2026-09-04 • Version 1.1
Catalogue review recorded · individual claims have separate datesAnalysisMethodology
AI Bench Index was built to counteract the proliferation of programmatic AI content and unverified benchmark aggregation. We do not invent benchmark scores, we do not fabricate consensus averages, and we do not create programmatic SEO comparison pages. Every data point published on this platform must earn its place through verifiable provenance.
How to read an evidence record
The source answers “who reported this?”; the run conditions answer “what was actually tested?”. Neither is replaced by a recent page-update date.
- Ind · independent evaluator
- Reported outside the model’s provider. Independence is not proof that the test is representative or reproducible.
- Own · benchmark owner
- Reported by the organisation maintaining that evaluation. A benchmark owner may also have commercial interests.
- Prov · model provider
- The model developer’s report. Preserve its configuration and limitations; do not present it as our measurement.
- Our measurement · first-party receipt
- Requires the actual host, model/runtime version, settings, command and output artefacts. Example fixtures never qualify.
A fairer comparison, not a universal rank
Receipt-based ranks and charts separate benchmark/version, unit, evaluator, configuration, recorded inference settings and evidence state. Missing or unevaluated conditions stay unranked. Matching fields are necessary, not sufficient: unrecorded prompt differences or uncertainty can still matter. Filtering does not renumber a result into a better rank.
Publisher boards are transcriptions of their published results, not independently reproduced experiments. Their own mixture of models and settings remains visible. Historical results from different scaffolds are not ranked against current ones. Small differences without uncertainty estimates are descriptive, not statistically established superiority.
AI assistance and editorial responsibility
AI assists research, drafting, data organisation and software development. It is not a source or a tester. Source links, test receipts and explicit unknowns carry the evidence. The site owner remains responsible for approval before publication; a source-review label does not imply an independent human replication.
Corrections and changing information
A correction should identify the page and entry, the disputed value, a primary source and its date. Correct the canonical record, preserve the previous observation when it is historically valid, record what changed and check every page derived from it. Recheck commercial terms and prices more frequently than architecture explanations. Never update all verification dates because a page was redesigned.
This revision’s correction record · 4 September 2026
- Arena Agent baseline: average model, not human performance. Source: Arena methodology.
- SWE-bench Verified: preserve historical results and attribute the frontier-evaluation critique. Source: OpenAI’s February 2026 audit.
- CodeRabbit: current trial and hourly limits use the current plan names; self-hosting is Enterprise-only. Source: CodeRabbit plans.
- Hardware: community and experimental support are no longer collapsed into an unconditional “yes”. Memory-fit predictions were removed from the information catalogues.
1. Data Provenance & Source Hierarchy
Every benchmark score and specification on AI Bench Index is treated as an observation rather than an absolute property of a model. Sources are classified by their relationship to the claim, not by a blanket guarantee of accuracy:
- Provider reports — what the developer publishes: Research papers (not necessarily peer-reviewed), system cards, official technical reports, and model cards authored directly by the model creators (e.g. OpenAI System Cards, Anthropic Model Cards, Google DeepMind Reports, Meta AI papers).
- Independent evaluator / benchmark owner: Official evaluation runs executed and hosted by independent academic labs and consortiums (e.g. TIGER Lab MMLU-Pro, Princeton NLP SWE-bench Verified leaderboard, ARC Evals).
- Community / third-party reports: Replications performed by reputable third parties using published, public harnesses with deterministic random seeds.
2. Treatment of Conflicting Benchmark Results
Frontier models frequently receive differing scores on the same benchmark due to differences in prompt harnesses, sampling temperatures, system instructions, few-shot samples, and scaffolding (e.g. single-turn pass vs agentic loop).
Rather than flattening conflicting scores into an arbitrary single average (such as model.score = 75), AI Bench Index preserves distinct observation records. Each observation records:
3. Date Semantics & Freshness Integrity
We enforce strict date semantics across all entities. A page rebuild date is never passed off as data freshness. We track:
4. Benchmark Versioning & Contamination Monitoring
Over time, popular benchmark datasets suffer from data contamination (when evaluation questions leak into model pretraining corpora) and question errors.
AI Bench Index distinguishes benchmark versions rigorously. For example, results on SWE-bench Verified (500 human-validated problem statements with verified tests) are tracked separately from original SWE-bench. Older benchmark versions are tagged as superseded but preserved for historical analysis.
5. Pricing Verification & History
Pricing is multi-dimensional. We do not reduce pricing to a single number. We separate input tokens, cached prompt tokens, output tokens, batch execution discounts, and context-tiered rates. When providers update prices, we do not overwrite historical rates; we close the effective date window (effectiveTo) and append the new rate structure.
6. Current Coverage & Independence Notice
Current coverage and independence:AI Bench Index catalogues vendor-published, benchmark-publisher, third-party and first-party records as distinct evidence classes. Unless a record is explicitly labelled as an AI Bench Index run with a downloadable receipt, the site has not reproduced the underlying evaluation. Transcription checks confirm what a cited source reported; they do not independently validate the publisher's experiment.
Independence: AI Bench Index does not accept paid rankings, sponsored placements, or preferential scoring from any AI model lab.
7. First-party local inference laboratory
Local AI pages estimate memory headroom against a stated threshold and assumptions; a “fits” label means the estimate clears that threshold, not a guarantee of usable performance. They do not invent tokens/sec. First-party speed numbers are published only from a validated receipt that names the exact model artifact, runtime version, host, and versioned profile. Parser fixtures (including CUDA llama-bench samples from runtime docs) are never shown as AIBenchIndex measurements.
- Profiles: local-generation-v1, local-prompt-processing-v1, local-long-context-v1, local-end-to-end-v1, local-serving-v1. Methodology changes mint v2.
- Classes stay separate: microbenchmark ≠ end-to-end ≠ server throughput.
- Published throughput is the median of warmed repetitions (llama-bench default: 5).
- Estimated vs measured memory is recorded; formulas are not auto-tuned from a handful of runs.
- No
/local-ai/benchmarksindexable listing until the dataset is substantial.
Full lab notes are published as part of the Local AI section alongside each first-party receipt, including the exact model artifact, runtime version, host, and versioned profile.
