Skip to main content

What it can do

  • Review a whole benchmark and flag what is likely wrong, with the evidence
  • Catch duplicate questions and an expected answer that is not among its own options
  • Catch correct answers piled onto one letter, and groups too small to mean anything
  • Catch runs scored under earlier files, whose scores are no longer comparable
  • Catch a leaderboard column nothing writes to, an ended phase, per-question results switched off
  • Read one run question by question: expected, answered, right or wrong
  • Show only the wrong answers
  • Say what the scoring found missing on each miss, instead of guessing why
  • Compare two runs: what got fixed and what broke
  • Read every question across recent runs: always missed, always right, and the ones that separate models
  • Decide whether a miss is the benchmark’s fault or the model’s
  • Score a run’s saved answers again under a changed scoring, with no model call, to prove a fix
  • Propose the fixes in order and hand the chosen ones over

Prompts to try

Why an answer was marked wrong

The reason comes from what the scoring saw, never from reading the answer:
Q5 expected the key points “digest”, “intestine” and “gut”. The model wrote “intestinal twists”. The scoring matches whole words, so “intestinal” does not contain “intestine”.
That is what separates a benchmark problem from a model problem. It is likely the benchmark when every model misses the same question, when the answer is right in words the key does not accept, or when the question can be read two ways. It is likely the model when other models get it right, or the answer is simply wrong rather than differently worded. Training a model against a broken answer key teaches it the wrong answer, which is why the sorting happens before anything is proposed.

Proving a fix without paying for another run

The change is made in a local copy and the run’s already saved answers are scored again. No model is called, so the test costs nothing. What comes back is a count, not an opinion: how many questions turn right, and how many break. A fix that repairs four answers while breaking two is not a fix.
The BenchGen AI agent proving a scoring fix on a run's saved answers before applying it: before the fix 3 of 10 right, after the widened key 7 of 10 right, four questions now right, three still wrong for stated reasons, and zero newly wrong

A fix proved on a run's saved answers before it was applied: 3 of 10 to 7 of 10, none broken

On the run above, applying the proved fix and running the same model again scored 6 out of 10, with no question made easier. The full story is in the guide.

What it will not do

  • Change anything. It proposes, asks, and hands over
  • Guess why an answer was marked wrong, or report a score the tools did not produce
  • Treat a re-score as a result. Those are old answers under changed rules, never a leaderboard entry
Benchmarks built before the platform recorded per-question results cannot be read question by question yet. The review says so and offers the fix, and only runs made afterwards carry the detail.

Edit a benchmark

Apply what the analysis proposed.

Train a model

When the misses really are the model’s own.

Per-question results (API)

One run, question by question, from your own code.

Question statistics (API)

Every question across recent runs.
Last modified on September 24, 2026