Back to the journal
Research / Analysis

An AI Test Is Not Ready Until Its Grader Has Been Tested

An agent can produce a convincing evaluation while omitting basic checks of the scoring system. Validating the grader should be part of the deliverable.

An empty frame and an orange test block sit beside a scale producing a blank results scroll.

AI agents can generate tests without establishing that those tests grade answers correctly. In a study updated September 10, AIMultiple reports that none of 32 benchmark submissions from 16 models demonstrated either a blank-answer check or a correct-answer check of its scorer. The tasks covered converting questions into database queries and selecting tools with their arguments.

The failure is especially revealing because an evaluation is supposed to create confidence in something else. If its own scoring rules are unexamined, it can manufacture confidence instead. A table of results may look more authoritative than an ordinary model answer while resting on an equally untested assumption. The lesson is to make the grader’s behavior an explicit part of what an agent must demonstrate.

A blank-answer check asks whether an empty response earns almost nothing. A correct-answer check asks whether the expected answer survives the full grading process and receives full credit. Without those controls, a polished results table can leave a basic question unanswered: does its measuring instrument behave as intended?

The instrument needs a check

A known wrong answer and a known right answer establish basic reference points. They do not prove that every judgment between those points is sensible. But a grader that cannot distinguish them has not earned the authority to rank more ambiguous responses. These checks are modest precisely because they concern the foundation. Sophisticated analysis built on top of an unreliable foundation cannot repair it.

There is an important distinction between a program that runs and a measurement that means something. A scoring script can execute without errors, write a results file and produce a neat chart. None of those outputs establishes that its values track the quality the evaluator intended to measure. The visible completeness of an automated workflow can make that distinction harder to notice unless someone asks for the underlying demonstration.

The main prompts left validity checks to the agents. One separate pilot explicitly requested the controls and produced them, but the authors say that prompt’s effect across models remains untested. Different agent programs, retained attempts and criteria graded by model judges also complicate a simple comparison of underlying models.

Omission is a different finding from inability

The separate pilot changes how the result should be interpreted. It suggests that asking explicitly for the controls can matter in at least that instance. It does not show that every model would respond in the same way, or that wording explains all of the omissions. The justified criticism concerns what the retained submissions demonstrated under the tested instructions. A broader claim of incapacity would outrun the experiment.

That narrower interpretation still has practical force. If a task is important enough to delegate, it is worth specifying what would count as a valid deliverable. A request to build a benchmark should include evidence that the grader behaves sensibly on cases whose outcomes are known. Leaving that requirement implicit and then treating a polished output as complete makes the user’s confidence depend on an assumption the workflow never checked.

The study comes from a firm that also sells AI automation services. Agentive has not independently reproduced it. Its findings support scrutiny of these submitted benchmarks; they do not establish that every AI-authored benchmark is invalid or that agents are incapable of writing useful tests.

The study’s commercial context is another reason to read the method closely, not a reason to dismiss the result automatically. A provider can publish useful evidence about a problem it also sells services to address. Readers need enough detail to assess the evidence without adopting the provider’s sales conclusion. Here, the limited setup supports a concern about these submissions; it cannot establish a universal ordering of agent reliability.

Broader evaluation research supplies a relevant standard of transparency. Stanford’s HELM project released raw prompts and model completions alongside a toolkit for examining results. That earlier work is independent background, not confirmation of AIMultiple’s experiment. It illustrates why a published score is easier to assess when the underlying evaluation is available for inspection.

Transparency helps only when someone can use it

Making prompts and responses available creates an opportunity to inspect the path from task to conclusion. It does not ensure that every reader will do so. A good evaluation report should therefore explain the decisive checks plainly and make their results easy to locate. Transparency is most useful when it reduces the effort required to challenge the conclusion, rather than burying that opportunity in an enormous archive.

A stronger acceptance process would start with the known cases, then consider plausible partial answers and failures that the task is meant to detect. The choice of those cases requires judgment about what the evaluation is for. Passing the easiest controls is necessary evidence, not a guarantee that the full benchmark captures the desired capability. The aim is to expose assumptions progressively instead of allowing an aggregate score to conceal them.

The conclusion should change what buyers of automated evaluations ask to receive. A benchmark package should include the test, the grader and evidence that the grader distinguishes relevant outcomes. A ranking can come afterward. Until that foundation is visible, the result is an evaluation proposal, however polished its charts may be. Agents can help build the measuring instrument; their ability to generate it does not exempt it from being measured.

Explore More Stories ↗