> ## Documentation Index
> Fetch the complete documentation index at: https://benchgen.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# What the agent is allowed to do

> The BenchGen AI agent acts as you, with a credential scoped to one chat turn. Here is what it can change, what it always asks first, and what it can never do.

The agent acts **as you**. It sees your benchmarks, your models and your runs, and nobody else's.
Everything it creates belongs to you, and everything it spends comes from your credits.

## How its access works

Each chat turn gets its own credential, minted for you and expiring with the turn. The agent never
sees your password, and it never sees a token you created yourself.

<Warning>
  You never paste an API token into the chat, and the agent never prints one or asks for one. A
  token in a chat message is a leaked token. If the agent ever asks for one, it is wrong: say no
  and tell us.
</Warning>

That credential carries only the permissions the platform grants it. Reading benchmarks and runs,
launching runs, creating and editing benchmarks and launching training are separate permissions,
and a deployment can switch the creating, editing and publishing ones off entirely.

## What it always asks before doing

The agent shows you exactly what it would do, then stops. Your next message decides.

| Action | What you see first |
| - | - |
| Create a benchmark | The title, the format, the question count and the questions with their answers |
| Publish a benchmark | Which benchmark, and that publishing makes it visible to everyone |
| Change a benchmark | The change, checked against the benchmark before anything is applied |
| Run a model | The benchmark and the model it picked, when your message left a choice open |
| Train a model | The full plan: job type, base model, dataset, hyperparameters and GPU availability |

Anything free-form, like a title or a question count, it asks in plain words. A choice between a
few options comes as buttons.

## What it can never do

These are not switched off by policy, they are outside its access entirely:

* Delete or unpublish a benchmark.
* Delete a model, a dataset or a training job.
* Edit a run's score by hand.
* Add or remove phases or tasks.
* Change who may edit or join a benchmark, including collaborators and whitelists.
* Set up the AI judge, which needs an API key.
* Email a benchmark's participants.
* Touch billing, or create, read and revoke API tokens.
* Add a finished run to a leaderboard. That is done from the run's page in the web app.

Ask for one of these and the agent says it cannot and points you at the right page in the app.

## What it does with your data

The agent reads what it needs for the task at hand: a benchmark's questions and answer key, a
run's per-question results, your model catalogue. When it works on a custom format, it downloads a
copy of the benchmark to its own scratch space, changes that copy, and nothing reaches the platform
until you approve the change.

Scoring a run's saved answers again, to prove a fix before applying it, happens entirely on that
local copy. No model is called, nothing is uploaded, and no leaderboard moves.

## Limits worth knowing

| Limit | What it means |
| - | - |
| A chat turn is capped at about 15 minutes | Very large benchmarks are written in parts. Reply `continue` to go on. |
| The agent cannot message you first | Ask "is it done?" rather than waiting for a run or a job to announce itself. |
| Benchmark creation is rate limited | 10 benchmarks a day and 3 a minute through the agent and the API. The web app is not limited. |
| Questions freeze after the first run | For question and answer and multiple choice benchmarks. Custom formats stay editable. |

## Related

<CardGroup cols={2}>
  <Card title="What the agent can do" icon="wand-magic-sparkles" href="/docs/skills/overview">
    The capabilities, one page each.
  </Card>

  <Card title="API tokens" icon="key" href="/docs/api-reference/introduction">
    The same permissions, when you call the API from your own code.
  </Card>
</CardGroup>
