> ## Documentation Index
> Fetch the complete documentation index at: https://benchgen.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Run a benchmark on a model

> The skill that browses benchmarks and models, launches a run on your account, and reports its state and score with a link to the run.

| The skill | `benchmark-launch` |
| - | - |
| **Changes anything** | Yes. A run spends your credits, exactly as a launch from the Evaluate page does. |
| **It picks this up when you say** | "what benchmarks can I run", "run this on Claude", "is my run done", "show my last runs" |
| **It hands over to** | `benchmark-analyze` to look inside a finished run |

## What it can do

* List the benchmarks you can run
* List the models you can run them on, both your own and the shared platform models
* Tell you which models are live right now and which have to be brought up
* Join a benchmark and queue a run on your account
* Report a run's state, and its score once it finishes
* List your recent runs with their scores and links
* Narrow that list to one benchmark or one model

## Prompts to try

```text theme={null}
What benchmarks can I run, and which of my models are live right now?
```

```text theme={null}
Run benchmark <id> on <model name>.
```

```text theme={null}
Is run <id> done? What did it score?
```

```text theme={null}
Show my last 5 runs with their scores and links.
```

## Run states

| State | What it means |
| - | - |
| `queued` | Waiting for capacity |
| `preparing` | The environment and the model are being brought up |
| `running` | The model is answering |
| `scoring` | The answers are being scored |
| `finished` | Done, with a score |
| `failed` | Stopped with an error, which the skill reports |
| `cancelled` | Stopped before it finished |

## What it will not do

* Launch without both a benchmark and a model. If either is unclear it shows the candidates and asks
* Launch twice for one request, or retry a failed launch on its own
* Wait for a run. A chat turn times out after about a minute
* Add a finished run to the leaderboard. That is done from the run's page in the web app

<Note>
  The agent cannot message you first. Ask "is it done?" rather than waiting to be told.
</Note>

## Related

<CardGroup cols={2}>
  <Card title="Analyze the results" icon="magnifying-glass-chart" href="/docs/skills/analyze-results">
    Why each answer was marked wrong, and whose fault it is.
  </Card>

  <Card title="Looking inside a run" icon="robot" href="/docs/guides/benchgen-ai-agent/working-with-the-agent#look-inside-a-run">
    The walkthrough, including what the model actually answered.
  </Card>

  <Card title="Run a benchmark (API)" icon="code" href="/docs/api-reference/endpoint/run-benchmark">
    Launch runs from your own code.
  </Card>

  <Card title="Run status (API)" icon="signal" href="/docs/api-reference/endpoint/run-status">
    Poll a run programmatically.
  </Card>
</CardGroup>
