What it can do
- List the benchmarks you can run
- List the models you can run them on, both your own and the shared platform models
- Tell you which models are live right now and which have to be brought up
- Join a benchmark and queue a run on your account
- Report a run’s state, and its score once it finishes
- List your recent runs with their scores and links
- Narrow that list to one benchmark or one model
Prompts to try
Run states
What it will not do
- Launch without both a benchmark and a model. If either is unclear it shows the candidates and asks
- Launch twice for one request, or retry a failed launch on its own
- Wait for a run. A chat turn times out after about a minute
- Add a finished run to the leaderboard. That is done from the run’s page in the web app
The agent cannot message you first. Ask “is it done?” rather than waiting to be told.
Related
Analyze the results
Why each answer was marked wrong, and whose fault it is.
Looking inside a run
The walkthrough, including what the model actually answered.
Run a benchmark (API)
Launch runs from your own code.
Run status (API)
Poll a run programmatically.