How its access works
Each chat turn gets its own credential, minted for you and expiring with the turn. The agent never sees your password, and it never sees a token you created yourself. That credential carries only the permissions the platform grants it. Reading benchmarks and runs, launching runs, creating and editing benchmarks and launching training are separate permissions, and a deployment can switch the creating, editing and publishing ones off entirely.What it always asks before doing
The agent shows you exactly what it would do, then stops. Your next message decides.
Anything free-form, like a title or a question count, it asks in plain words. A choice between a
few options comes as buttons.
What it can never do
These are not switched off by policy, they are outside its access entirely:- Delete or unpublish a benchmark.
- Delete a model, a dataset or a training job.
- Edit a run’s score by hand.
- Add or remove phases or tasks.
- Change who may edit or join a benchmark, including collaborators and whitelists.
- Set up the AI judge, which needs an API key.
- Email a benchmark’s participants.
- Touch billing, or create, read and revoke API tokens.
- Add a finished run to a leaderboard. That is done from the run’s page in the web app.
What it does with your data
The agent reads what it needs for the task at hand: a benchmark’s questions and answer key, a run’s per-question results, your model catalogue. When it works on a custom format, it downloads a copy of the benchmark to its own scratch space, changes that copy, and nothing reaches the platform until you approve the change. Scoring a run’s saved answers again, to prove a fix before applying it, happens entirely on that local copy. No model is called, nothing is uploaded, and no leaderboard moves.Limits worth knowing
Related
What the agent can do
The capabilities, one page each.
API tokens
The same permissions, when you call the API from your own code.