Skip to main content
The agent acts as you. It sees your benchmarks, your models and your runs, and nobody else’s. Everything it creates belongs to you, and everything it spends comes from your credits.

How its access works

Each chat turn gets its own credential, minted for you and expiring with the turn. The agent never sees your password, and it never sees a token you created yourself.
You never paste an API token into the chat, and the agent never prints one or asks for one. A token in a chat message is a leaked token. If the agent ever asks for one, it is wrong: say no and tell us.
That credential carries only the permissions the platform grants it. Reading benchmarks and runs, launching runs, creating and editing benchmarks and launching training are separate permissions, and a deployment can switch the creating, editing and publishing ones off entirely.

What it always asks before doing

The agent shows you exactly what it would do, then stops. Your next message decides. Anything free-form, like a title or a question count, it asks in plain words. A choice between a few options comes as buttons.

What it can never do

These are not switched off by policy, they are outside its access entirely:
  • Delete or unpublish a benchmark.
  • Delete a model, a dataset or a training job.
  • Edit a run’s score by hand.
  • Add or remove phases or tasks.
  • Change who may edit or join a benchmark, including collaborators and whitelists.
  • Set up the AI judge, which needs an API key.
  • Email a benchmark’s participants.
  • Touch billing, or create, read and revoke API tokens.
  • Add a finished run to a leaderboard. That is done from the run’s page in the web app.
Ask for one of these and the agent says it cannot and points you at the right page in the app.

What it does with your data

The agent reads what it needs for the task at hand: a benchmark’s questions and answer key, a run’s per-question results, your model catalogue. When it works on a custom format, it downloads a copy of the benchmark to its own scratch space, changes that copy, and nothing reaches the platform until you approve the change. Scoring a run’s saved answers again, to prove a fix before applying it, happens entirely on that local copy. No model is called, nothing is uploaded, and no leaderboard moves.

Limits worth knowing

What the agent can do

The capabilities, one page each.

API tokens

The same permissions, when you call the API from your own code.
Last modified on September 24, 2026