Vendor AI Benchmarks: Only Numbers With Logs Belong in the Deck

A vendor-run benchmark is a screening signal. Business cases should cite independent runs with per-task logs, led by Epoch AI and Artificial Analysis, plus a 50-task replication on your own data.

By Rajesh Beri·October 7, 2026·16 min read
Share:
A procurement meeting table with a printed vendor slide showing a bar chart, a red pen beside it, and an open laptop displaying a long spreadsheet of individual test results with pass and fail marks.

Illustration generated using AI

A benchmark score the vendor ran itself, on a task set you can't see, at settings it chose, belongs in the screening pile and nowhere near your business case. The numbers that survive a CFO's questions come from two places: an independent evaluator that publishes its settings and per-task logs (Epoch AI's Benchmarking Hub first, Artificial Analysis second), and a 50-task replication you run on your own data at the configuration you will actually pay for. Use private-set evaluators like Vals AI and Scale's leaderboards as a contamination check. Leave Arena (formerly LMArena) rankings out of procurement documents entirely.

The gap between a vendor number and an independent rerun is routine. Anthropic reported 65% for Claude 3.5 Sonnet on GPQA Diamond; Epoch AI's 16-run mean was 55%, a difference Epoch attributes to evaluation settings. Neither number was dishonest. Only one of them is reproducible by someone outside the lab.

Evidence source Who runs it Task set public? Per-task logs? Contamination control Conflict to weigh Price (checked Oct 7, 2026) Business-case use
Vendor model card / launch post The vendor Usually, for public benchmarks Rarely Vendor's word Total Free Screening only
Epoch AI Benchmarking Hub Epoch, on Inspect Yes for most; FrontierMath mostly private Yes, per question Holdout sets on FrontierMath OpenAI funded FrontierMath Free, data CC BY 4.0 Cite it
Artificial Analysis AA, own harness and prompts Mixed; several private sets No public per-task logs Private held-out splits Sells private benchmarking to labs Leaderboards free to read; Pro $417/seat/month Cite with settings
Vals AI Vals, with domain institutions No, private by design No Strong (private) Undisclosed commercial terms Contact sales Cross-check
Scale Labs leaderboards Scale, humans plus LLM judges Mixed; private sets No Private sets, first-exposure rule Sells training data; Meta owns 49% Free to read Cross-check
Arena Crowd votes Prompts are user-submitted No None in the usual sense Labs pay for pre-launch testing Free to read Leave out

What Separates a Vendor-Run Benchmark From an Independent One?

A benchmark is independent when someone other than the model's maker chose the prompts, the decoding settings and the scoring, and published enough that a third party could rerun it. Ownership of the task set matters less than who controlled the run.

Three signals tell you which kind you are holding. First, look for a methodology page with temperature, reasoning effort, attempts per task and the harness named. Artificial Analysis publishes all of these: zero-shot prompting, temperature 0 for non-reasoning models and 0.6 for reasoning models unless the lab recommends otherwise, and between one and five repeats depending on the eval. A vendor launch chart almost never states the effort setting, which is how a model ends up benchmarked at top effort while the price you are quoted assumes the default.

Second, check whether the score can be traced to individual tasks. Epoch exposes per-question logs for the benchmarks it runs itself, including GPQA Diamond and SWE-bench Verified, through a log viewer on its hub. That is the only way to find out that a "win" came from 12 tasks the grader scored generously.

Third, read the funding disclosure, and read it skeptically when it is missing. Epoch is the most transparent evaluator in this table and still failed to disclose that OpenAI funded FrontierMath until o3's announcement. Its January 2025 post says OpenAI has the problems and solutions except for a 50-problem holdout set. Contributors who wrote the problems were not told either. Epoch fixed it by publishing the arrangement, which is why it still tops the list: you can now see the conflict and route around it by using the holdout numbers.

Is the Task Set Published, and Can You Reproduce the Run?

You can reproduce a result only if three things are public: the tasks, the harness, and the exact configuration. Most vendor claims publish the first and omit the other two.

The harness matters more than buyers expect. Swapping the scaffold around the same model can reorder a coding leaderboard, and Artificial Analysis states that its reimplementations of some outside benchmarks are "not directly comparable" to the original authors' published results, per its own methodology notes. If the evaluator that does this full time warns you that two runs of the "same" benchmark disagree, a vendor's chart against a competitor's self-reported number is two different experiments on one slide.

Private sets trade reproducibility for contamination resistance. Vals AI says its "scores are based on our privately held test sets to preserve the integrity and signal of our results." Scale's leaderboards combine private sets with open ones and feature a model only "the FIRST TIME when an organization encounters the prompts." Both designs make it harder for a lab to train on the test. Both also mean you cannot check a single answer. When a vendor cites a private-set win, as OpenAI did with Astra for Law's 54% on a Vals set nobody outside can read, treat it as a directional cross-check only.

How Do You Spot Contamination and Train/Test Overlap?

Contamination is when benchmark questions, or their answers, appeared in a model's training data, so the score measures recall of the test instead of the skill the test was built to measure. You usually can't detect it from outside, which is why the burden belongs on the vendor.

The evidence that it happens at scale is no longer contested. Scale researchers built GSM1k, about 1,250 new grade-school math problems written without any LLM help to mirror GSM8K, and found accuracy drops of up to 13% in the first version of the paper, with Phi and Mistral families showing systematic overfitting. The frontier models of that period showed little sign of it. In February 2026 OpenAI said it would stop reporting SWE-bench Verified because gains there increasingly reflected how much a model had seen the benchmark in training, and it recommended SWE-bench Pro instead. When a lab retires the coding benchmark every vendor deck was citing, any SWE-bench Verified number in a 2026 proposal deserves a question.

Disclosure is thin. A Stanford-led position paper reviewed 30 model developers and found only 9 report train-test overlap; the authors argue that outsiders cannot compute overlap without access to the training corpus.

Ask the vendor three questions in writing:

  1. What train-test overlap did you measure for each benchmark you cite, and by what method (n-gram match, embedding similarity, canary strings)?
  2. Was any benchmark data, or data derived from it, used in post-training or in selecting checkpoints?
  3. Which of the cited scores come from a held-out or post-cutoff split?

A vendor that answers "we follow industry best practices" has told you it did not measure. Arena adds its own version of the problem: The Leaderboard Illusion found Meta tested 27 private variants before the Llama 4 release, and estimated that Google and OpenAI received 19.2% and 20.4% of all Arena data. The paper estimates that access to Arena data can produce relative gains of up to 112% on the Arena distribution, which is overfitting to the voters.

What Does "Matches or Exceeds" Actually Permit?

"Matches or exceeds" lets a vendor claim parity on a benchmark it lost, as long as the loss sits inside a margin it did not state, on a subset of benchmarks it picked, at settings it chose. Each of those is a degree of freedom, and together they make the phrase nearly impossible to falsify.

Run the arithmetic. On a 500-task benchmark where both models score around 70%, the 95% interval on each score is roughly ±4 points (the standard error of a proportion is the square root of p times 1 minus p over n). A model 3 points behind can "match" honestly. Evan Miller's paper on error bars for evals lays out the formulas for comparing two models on the same questions; few launch charts use them. Our own reanalysis found only 3 of 36 reported model gaps survived a confidence interval.

Then count the denominator. A release that "beats" a rival can do it on one of 14 shared benchmarks and lead with that one. Ask for the full list of benchmarks the vendor ran, including the ones it didn't publish.

Then check the variant. In April 2025 Meta's Arena entry was an "experimental chat version" optimized for conversationality; the unmodified release ranked 32nd when Arena tested it. Your contract should name the model version and configuration the benchmark used, and the one you are buying should match it.

On the legal side, the ad self-regulator is watching AI claims but has not yet ruled on a model-versus-model benchmark comparison that we could find. Advertising lawyers at Faegre Drinker advise building "product-specific evidentiary records" before making AI claims, so a vendor that followed that advice has a record it can hand you on request.


How Each Evidence Source Holds Up, and Who Should Not Rely on It

Every source here has a buyer it fails. The ranking above reflects one test: could you defend this number to a skeptical finance lead with the evidence in hand?

Vendor Model Cards and Launch Posts

The vendor's own numbers are useful for one job: deciding whether a model is worth testing at all. They are the only source on day one, and labs publish more detail than they used to. Nobody should put them in a business case unchanged, because every degree of freedom described above was exercised by the party that profits from the result. Treat an unpublished score attached to a latency claim as no score at all.

Epoch AI Benchmarking Hub

Epoch is the source to cite first. It runs GPQA Diamond, MATH Level 5, Mock AIME, FrontierMath and SWE-bench Verified itself on the open-source Inspect framework, publishes per-question logs for most of them, licenses its data CC BY 4.0, and tracks 88 benchmarks in total, citing the source of each external data point. It costs nothing.

Who should not lean on it: a buyer choosing a model for an enterprise workflow such as claims triage or contract review. Epoch's self-run set is math, science and coding heavy, and its external entries are only as good as the leaderboard they came from. Weigh its FrontierMath figures for OpenAI models with the funding arrangement in mind and prefer the holdout results.

Artificial Analysis

Artificial Analysis is the best single cross-check for comparing many models at once. Its Intelligence Index v4.3.2 combines 10 evaluations across agents, coding, general and scientific reasoning, several of them on private held-out splits, and it estimates a 95% interval of less than ±1% on the overall index. It also reports cost to run, which vendors rarely volunteer. The public leaderboards are free to read; Pro is $417 per seat per month and Enterprise is custom, checked October 7, 2026.

Who should not lean on it: anyone who needs to audit an individual answer, since there are no public per-task logs. Some of its evals are graded by a single LLM judge (GPT-5.6 Luna on AA-Omniscience, AA-LCR and HLE), and others by judge panels; check which before citing a narrow gap between an OpenAI model and a rival. It also sells private benchmarking to the labs it ranks, which its founders say does not touch the public tables. Because it is disclosed, you can account for it.

Vals AI

Vals is the right cross-check for legal, tax and finance work, because it builds tasks with domain institutions and keeps the test sets private. That design resists contamination better than anything public. Pricing is contact-sales only; its site lists no plans.

Who should not lean on it: a buyer who needs a number to survive an audit. You cannot inspect the tasks, the rubric or a single graded response, and its commercial relationships with the labs it scores are not published. Use a Vals score to build a shortlist, then settle the purchase with your own replication. We track it on our Vals AI tool page.

Scale Labs Leaderboards

Scale's leaderboards pair private datasets with open ones and let a model appear only on its first exposure to the prompts, a sensible guard against labs tuning to the test. Humans design the evaluations and LLMs scale them. Reading them is free.

Who should not lean on it: anyone comparing Meta's models with a rival's. Meta took a 49% non-voting stake in Scale for $14.3 billion in June 2025, and Scale's core business is selling training data to the labs whose models it ranks. See Scale AI on our tools directory.

Arena (formerly LMArena): The Loser for Procurement

Arena measures which answer anonymous users preferred in a side-by-side vote. That is a real signal for chat tone and formatting. It is the wrong signal for a procurement document, and it loses this comparison for three reasons on the record: labs can test many private variants and publish the best one (27 for Meta before Llama 4); the largest labs get far more battle data than open-weight entrants; and the company, which raised a $150 million Series A in January 2026, sells pre-launch evaluation to the same labs, according to reporting on the round. A preference rank tells you nothing about whether a model will extract the right clause from your contracts.

Who should use it anyway: a team picking a model for a consumer-facing chat surface where style is the product. Everyone else should leave it out of the deck.


How to Run a 50-Task Replication on Your Own Data

A 50-task replication is a small, blinded test of the vendor's claim on your own work, at the configuration you will buy, scored against answers you wrote down before running anything. It takes about a week and it is the only evidence in this guide that measures your problem.

  1. Pick 50 tasks from your own backlog that have never been on the public internet: closed tickets, redacted contracts, last quarter's escalations. Public examples risk the contamination you are trying to avoid. Hand-written beats synthetic, as our RAG eval dataset comparison found.
  2. Write the pass criterion for each task first, as a reference answer or a checklist a reviewer can apply in under a minute.
  3. Run the candidate and your incumbent at the tier you will pay for: same reasoning effort, same context length, same tools. Three runs each, so you see variance.
  4. Grade blind: strip model names before a human or an LLM judge sees outputs. If you use an LLM judge, use one from a third model family and spot-check 10 grades by hand.
  5. Report a pass rate with its interval. At 50 tasks and an 80% pass rate, the 95% interval is roughly ±11 points.

That last number sets what the test can and cannot do. Fifty tasks will catch a vendor whose 85% claim comes in at 60% on your data, which is the failure that causes regret. They will not resolve a 4-point gap between two good models; for that you need several hundred tasks. Say which question you answered.

Tooling is not the constraint. Promptfoo is MIT licensed and runs from a YAML file; it is now part of OpenAI, so if OpenAI is one of your candidates, note that in the write-up. Inspect, developed by the UK AI Security Institute and Meridian Labs, supports over 20 model providers and custom scorers, and it is what Epoch runs on. Hosted options start free: Braintrust's Starter plan is $0 with 10,000 scores and 14-day retention, and Langfuse's Hobby tier is free for 50,000 units a month, both checked October 7, 2026. See Braintrust and Langfuse in our directory, and our comparison of the three eval tools if you plan to keep the set as a regression gate.

Keep the 50 tasks after the purchase. Vendors change models behind a stable name, and an eval that never reruns won't notice.

Which Claims Belong in the Business Case?

Put a number in the business case only if you could hand a skeptic the source, the settings and the logs. That rule keeps four kinds of evidence and drops four.

Put these in:

  1. Your own 50-task pass rate, with its interval, the configuration and the date.
  2. Epoch results for the benchmarks closest to your workload, linked to the logs.
  3. Artificial Analysis index or per-eval scores, with the version number and the reasoning setting.
  4. Cost per completed task from your replication, not price per million tokens.

Leave these out:

  1. Any vendor-run score without stated effort, attempts and harness.
  2. "Matches or exceeds" claims, unless the vendor supplies intervals and the full benchmark list.
  3. Arena rank, for anything other than a consumer chat product.
  4. SWE-bench Verified numbers from 2026, given OpenAI's own retirement of the benchmark.

Private-set results from Vals or Scale go in an appendix as corroboration. If they disagree with your replication, believe the replication and say so in the document.

What Changes the Answer?

The criteria that predict regret are the ones a benchmark can't capture: whether the model holds up on your inputs, whether the version you tested is the version you get, and whether the cost per completed task matches the pitch.

Three situations change the recommendation above. If you are in a regulated domain with no public benchmark close to your work, Vals-style private evaluation plus a larger replication (200 tasks or more) beats every public leaderboard. If you are choosing among open-weight models you will host yourself, you control the configuration, so Epoch and Artificial Analysis numbers transfer better and the replication can be smaller. If the vendor offers to run the bake-off for you, decline; a vendor-run replication on your data inherits every problem of a vendor-run benchmark, with your logo on it.

This Week:

  1. Send the three contamination questions to every vendor on your shortlist and ask for the full list of benchmarks they ran.
  2. Pull 50 tasks from your own closed work and write the pass criteria before anyone runs a model.

This Month:

  1. Run the replication at the configuration in the quote, three runs per model, graded blind.
  2. Rebuild the business case's evidence section using only the four "put these in" sources.

Before Signature:

  1. Name the model version and configuration in the contract, and require notice before either changes.
  2. Schedule a rerun of the 50 tasks for 90 days after go-live.

The Bottom Line

Enterprise databases went through this in the 1990s, when vendor-run performance claims were common enough that the industry built audited benchmarks with full disclosure reports to restore trust. AI model evaluation is partway through the same correction: Epoch publishes logs, Artificial Analysis publishes settings, and one lab has publicly retired the benchmark that had become a training target. None of it has reached the vendor launch chart yet.

Until it does, the job falls to the buyer. Send the contamination questions and pull your 50 tasks this week.

Continue Reading

Share:

Frequently Asked Questions

How can I tell if an AI benchmark result is independent?

Check who chose the prompts, decoding settings and scoring, and whether they published enough to rerun it. Independent evaluators like Epoch AI and Artificial Analysis publish methodology pages with temperature, attempts and harness; Epoch also publishes per-question logs. A vendor launch chart usually states none of these.

What is benchmark contamination?

Contamination means benchmark questions or answers appeared in a model's training data, so the score measures recall of the test. Scale's GSM1k study found accuracy drops of up to 13% on fresh problems in the paper's first version, and OpenAI stopped reporting SWE-bench Verified in February 2026 over contamination.

Is LMArena a reliable source for choosing an enterprise model?

Not for procurement. Arena ranks crowd preference votes, and The Leaderboard Illusion paper found Meta tested 27 private variants before Llama 4's release. It is a reasonable signal for consumer chat tone, but it says nothing about accuracy on your own tasks.

How many examples do I need to check a vendor's benchmark claim?

About 50 tasks from your own data catches large gaps, such as an 85% claim landing at 60%. At an 80% pass rate the 95% interval on 50 tasks is roughly plus or minus 11 points, so resolving a small gap between two strong models needs several hundred tasks.

What does 'matches or exceeds' mean in a vendor benchmark claim?

It usually means the vendor's score sits within an unstated margin of error on benchmarks the vendor selected, at settings it chose. On a 500-task benchmark near 70%, each score carries roughly a 4-point interval, so a model 3 points behind can still be said to match.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →