> ## Documentation Index
> Fetch the complete documentation index at: https://benchgen.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Edit a benchmark from chat

> The skill that changes a benchmark in place: its questions, scoring, answer key, leaderboard columns, settings and dates, each checked first.

| The skill | `benchmark-edit` |
| - | - |
| **Changes anything** | Yes. It changes the benchmark in place, after a check and your confirmation. |
| **It picks this up when you say** | "rename it", "fix question 4", "correct the answer", "change the scoring", "add a metric", "rank by accuracy", "extend the deadline" |
| **It hands over to** | `benchmark-launch` to run the model again so scores are comparable |

## What it can do

* Rename a benchmark and rewrite its description, terms and contact details
* Edit its pages and tabs
* Fix a question or its expected answer, while the benchmark has never been run
* Replace the scoring of a custom format, at any time
* Replace the answer key, the prompts or the results page of a custom format, at any time
* Add a leaderboard column, retitle it, re-sort it or hide it
* Change which column ranks the leaderboard, showing the top of the board before and after
* Turn detailed results, auto-approve and auto-run on or off
* Change the docker image or the logo
* Rename a phase, move its dates, set submission and time limits
* Check every change against the real benchmark before writing anything
* List the runs scored under the previous files, so you can run those models again

<Note>
  Questions freeze after the first run, because the scores describe the questions that run saw.
  That rule covers those questions only. A custom format's scoring, answer key and prompts stay
  editable for the life of the benchmark, which is what makes the analyse and fix loop possible.
</Note>

## Prompts to try

```text theme={null}
Question 4's expected answer is wrong. It should be 1969, not 1968.
```

```text theme={null}
Widen the key points on questions 1, 5 and 9 so reasonable synonyms count.
```

```text theme={null}
The scoring only matches whole words, so "gradually" fails "gradual". Fix the matching itself
rather than patching each question.
```

```text theme={null}
Add a column for key point coverage to the leaderboard and rank by it.
```

```text theme={null}
Turn on detailed results, allow 5 runs per person per day, and extend the deadline to the
end of the year.
```

## What it will not do

* Publish, unpublish or delete a benchmark
* Add or remove phases and tasks
* Change collaborators or who may join
* Set up the AI judge, edit a run's score by hand, or email participants
* Re-score earlier runs. They keep their old scores until each model is run again
* Create a second benchmark to fix a first. Two copies share a title and split the leaderboard

<Tip>
  A metric with no leaderboard column is dropped on the way to the board, so adding a scoring type
  is always two changes together and the skill makes both. Columns are never deleted and their
  keys are never renamed, because existing runs point at them. Hiding is the reversible way to
  take one off the board.
</Tip>

## Related

<CardGroup cols={2}>
  <Card title="Analyze the results" icon="magnifying-glass-chart" href="/docs/skills/analyze-results">
    Find out what to fix, with proof, before you fix it.
  </Card>

  <Card title="Fixing a benchmark step by step" icon="robot" href="/docs/guides/benchgen-ai-agent/working-with-the-agent#analyze-a-run-and-fix-the-benchmark">
    The walkthrough, with a real before and after.
  </Card>

  <Card title="Edit a benchmark (API)" icon="code" href="/docs/api-reference/endpoint/edit-benchmark">
    The same edits from your own code.
  </Card>

  <Card title="Leaderboard columns (API)" icon="table-columns" href="/docs/api-reference/endpoint/change-leaderboard-columns">
    Add, retitle, re-sort and hide columns programmatically.
  </Card>
</CardGroup>
