Who Grades the AI Tests? Epoch Found Problems in Nine Major Benchmarks

Who Grades the AI Tests? Epoch Found Problems in Nine Major Benchmarks

Epoch AI’s benchmark reviews found incorrect answer keys, grading defects, and inconsistent evaluation methods across nine tests. Four others passed, with caveats that explain what their scores can actually tell us.

Written By
Corey Noles
Corey Noles
Sep 18, 2026
5 minute read

Before you trust an AI’s report card, someone should probably check the exam.

Epoch AI has been doing exactly that. Its benchmark review registry lists nine benchmarks as flawed, four as verified, and two with insufficient information to support a verdict, as of September 18. The flagged names include Humanity’s Last Exam, SWE-bench Verified, and the Berkeley Function Calling Leaderboard. Epoch’s benchmark registry puts those assessments alongside the tests used to compare AI systems.

The findings raise an uncomfortable question: When a model’s score changes, how much reflects its ability—and how much reflects the test?

Epoch’s reviews document cases where correct answers can lose points, incorrect answers can earn them, and changes in evaluation methods make leaderboard results difficult to compare. They also identify benchmarks that provide useful measurements, provided readers understand their limits.

That last part deserves attention. A “verified” badge still comes with reading material.

What Epoch actually checked

Under Epoch’s review framework, “Verified” means a benchmark’s results can broadly be interpreted as described. “Flawed” means substantive problems affect that interpretation. “Not enough information” means reviewers couldn’t obtain what they needed to make the call. Epoch separately labels its own benchmarks and excludes them from these reviews because of the conflict of interest.

The methodology looks at whether tasks and scoring rules can be inspected, whether grading works, whether results are comparable across versions, and whether the evaluation setup unfairly constrains or advantages models.

Its default scoring threshold for a flawed verdict is errors in at least 20% of the inspected sample, or a problem that corrupts grading at scale. Other failures, including mixing incomparable benchmark versions, can also trigger that designation.

Advertisement

Those percentages need careful handling. An error rate in a sample describes that sample. It doesn’t establish that the same percentage of every model’s score is wrong.

The nine benchmarks Epoch flagged

These verdicts apply to the versions or snapshots Epoch reviewed.


Some examples are almost painfully ordinary.

In BFCL, one task asks for the current time in Sydney and supplies a tool for retrieving a city’s local time. Epoch found that the answer key rewards declining to use a tool. Making the appropriate call earns zero. Another task relies on an outdated answer about a school’s headmaster, potentially penalizing a model for retrieving current information. Epoch’s BFCL review documents both.

Humanity’s Last Exam has its own answer-key problems. In one question, the author’s explanation leads to option E, while the recorded answer is D. In another, the answer key gives the reciprocal of the value calculated in its own rationale. Those are among the findings in Epoch’s HLE audit.

An AI can struggle because a question is difficult. It can also struggle because the exam is arguing with its own answer sheet.

There’s an especially interesting wrinkle for Neuron readers. We previously covered Datacurve’s DeepSWE and its critique of coding leaderboards. Epoch has now flagged DeepSWE v1.1 itself. Its reviewers used models to help surface potential errors, then checked the findings manually. They stopped after confirming enough false negatives to cross their threshold. The DeepSWE review therefore represents a partial audit.

Building a benchmark to address known evaluation problems still leaves plenty of room for new ones.


The four that passed, and what passing means

The verified group includes SimpleQA Verified, WeirdML v2, PostTrainBench v1.1, and ExploitBench v0.1. Each measures a different capability, and each has limitations.

SimpleQA Verified measures recall of obscure facts without tools. Epoch found score-affecting defects in five of 50 sampled questions—below its flawed threshold. It also notes that public questions and answers create contamination risk, while scoring abstentions as wrong means results partly reflect a model’s willingness to guess.

WeirdML v2 tests whether models can write PyTorch code for unfamiliar machine-learning problems under a fixed compute budget. Epoch praises its largely hidden tasks, novel datasets, and reporting of uncertainty. It also discloses that the review was written before it hired the benchmark’s creator.

Advertisement

PostTrainBench v1.1 provides useful evidence about models’ ability to fine-tune other models for specified evaluations. Its scope is narrower than autonomous AI research overall: improving one target score doesn’t establish that the resulting model gained broadly transferable abilities.

ExploitBench v0.1 measures progress toward exploiting known vulnerabilities in V8, a JavaScript engine. Epoch credits its graded measures of progress and verification approach, while flagging contamination concerns and uncertainty about how well results generalize to other cybersecurity work.

Taken together, these reviews show what verification can reasonably offer: a supported interpretation of a particular measurement, with its limitations attached.

What this changes for people choosing AI

For anyone comparing models, the practical implication is to read one layer below the score.

Check the benchmark version. Look at the kinds of errors the review found. Examine the tools, time limits, and other resources supplied during evaluation. Then test the shortlisted systems on representative work from your own operation.

A benchmark defect can push results in either direction. Loose grading can reward failure; overly strict grading can punish success. These reviews therefore don’t support a blanket conclusion that AI capabilities have been exaggerated by some fixed amount.

They do support more scrutiny of the instruments used to measure those capabilities.

Epoch’s process has limits, too. Its FAQ explains that reviewers stop once they find enough evidence for a flawed verdict, so the published problems may be incomplete. Fixing those problems doesn’t automatically earn verification. A new version needs another assessment, and Epoch’s capacity for repeat reviews is limited.

The useful outcome would be a stronger habit across the industry: publishing inspectable tests, maintaining clear versions, correcting grading errors, and explaining what a score actually supports.

The next time a model tops a leaderboard, ask what it accomplished—and whether the test would recognize a correct answer when it saw one.

Corey Noles

Corey Noles is the Host of The Neuron: AI Explained podcast and Managing Editor of AI and Experimental Content at TechnologyAdvice, where he leads the charge in testing and refining emerging content strategies across the company's portfolio.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.

Stay in the loop

Get notified when we publish new articles.