> ## Documentation Index
> Fetch the complete documentation index at: https://benchgen.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Analyze benchmark results

> The read-only skill that reviews a benchmark, reads a run question by question, and proves a scoring fix before proposing it.

| The skill | `benchmark-analyze` |
| - | - |
| **Changes anything** | No. Read only. It cannot edit a benchmark, launch a run or start training. |
| **It picks this up when you say** | "review my benchmark", "why did it score so low", "which questions failed", "compare these two runs", "how do I improve it" |
| **It hands over to** | `benchmark-edit` when the benchmark is at fault, `training-launch` when the model is |

## What it can do

* Review a whole benchmark and flag what is likely wrong, with the evidence
* Catch duplicate questions and an expected answer that is not among its own options
* Catch correct answers piled onto one letter, and groups too small to mean anything
* Catch runs scored under earlier files, whose scores are no longer comparable
* Catch a leaderboard column nothing writes to, an ended phase, per-question results switched off
* Read one run question by question: expected, answered, right or wrong
* Show only the wrong answers
* Say what the scoring found missing on each miss, instead of guessing why
* Compare two runs: what got fixed and what broke
* Read every question across recent runs: always missed, always right, and the ones that separate models
* Decide whether a miss is the benchmark's fault or the model's
* Score a run's saved answers again under a changed scoring, with no model call, to prove a fix
* Propose the fixes in order and hand the chosen ones over

## Prompts to try

```text theme={null}
Review benchmark <id> and tell me everything that looks wrong with it.
```

```text theme={null}
Show me only the wrong answers of run <id>.
```

```text theme={null}
Why was question 5 of run <id> marked wrong? Show me the answer in full.
```

```text theme={null}
Compare run <id> with the run before it: what got better and what got worse?
```

```text theme={null}
Analyse run <id>'s wrong answers, and prove any scoring fix you suggest before offering it.
```

## Why an answer was marked wrong

The reason comes from what the scoring saw, never from reading the answer:

> Q5 expected the key points "digest", "intestine" and "gut". The model wrote "intestinal twists".
> The scoring matches whole words, so "intestinal" does not contain "intestine".

That is what separates a benchmark problem from a model problem. It is likely the **benchmark**
when every model misses the same question, when the answer is right in words the key does not
accept, or when the question can be read two ways. It is likely the **model** when other models
get it right, or the answer is simply wrong rather than differently worded.

Training a model against a broken answer key teaches it the wrong answer, which is why the
sorting happens before anything is proposed.

## Proving a fix without paying for another run

The change is made in a local copy and the run's **already saved** answers are scored again. No
model is called, so the test costs nothing. What comes back is a count, not an opinion: how many
questions turn right, **and how many break**. A fix that repairs four answers while breaking two
is not a fix.

<Frame caption="A fix proved on a run's saved answers before it was applied: 3 of 10 to 7 of 10, none broken">
  <img src="https://mintcdn.com/benchgen-8fc81371/o4N0GNFEJXrQAuwu/images/guides/benchgen-ai-agent/08-proof-before-after.png?fit=max&auto=format&n=o4N0GNFEJXrQAuwu&q=85&s=e7c7a064b97620f1736c4b2cbb6755a3" alt="The BenchGen AI agent proving a scoring fix on a run's saved answers before applying it: before the fix 3 of 10 right, after the widened key 7 of 10 right, four questions now right, three still wrong for stated reasons, and zero newly wrong" width="960" height="245" data-path="images/guides/benchgen-ai-agent/08-proof-before-after.png" />
</Frame>

<Tip>
  On the run above, applying the proved fix and running the same model again scored 6 out of 10,
  with no question made easier.
  [The full story is in the guide](/docs/guides/benchgen-ai-agent/working-with-the-agent#analyze-a-run-and-fix-the-benchmark).
</Tip>

## What it will not do

* Change anything. It proposes, asks, and hands over
* Guess why an answer was marked wrong, or report a score the tools did not produce
* Treat a re-score as a result. Those are old answers under changed rules, never a leaderboard entry

<Warning>
  Benchmarks built before the platform recorded per-question results cannot be read question by
  question yet. The review says so and offers the fix, and only runs made afterwards carry the
  detail.
</Warning>

## Related

<CardGroup cols={2}>
  <Card title="Edit a benchmark" icon="pen-to-square" href="/docs/skills/edit-a-benchmark">
    Apply what the analysis proposed.
  </Card>

  <Card title="Train a model" icon="graduation-cap" href="/docs/skills/train-a-model">
    When the misses really are the model's own.
  </Card>

  <Card title="Per-question results (API)" icon="list-check" href="/docs/api-reference/endpoint/run-item-results">
    One run, question by question, from your own code.
  </Card>

  <Card title="Question statistics (API)" icon="chart-column" href="/docs/api-reference/endpoint/benchmark-item-stats">
    Every question across recent runs.
  </Card>
</CardGroup>
