What it can do
- Rename a benchmark and rewrite its description, terms and contact details
- Edit its pages and tabs
- Fix a question or its expected answer, while the benchmark has never been run
- Replace the scoring of a custom format, at any time
- Replace the answer key, the prompts or the results page of a custom format, at any time
- Add a leaderboard column, retitle it, re-sort it or hide it
- Change which column ranks the leaderboard, showing the top of the board before and after
- Turn detailed results, auto-approve and auto-run on or off
- Change the docker image or the logo
- Rename a phase, move its dates, set submission and time limits
- Check every change against the real benchmark before writing anything
- List the runs scored under the previous files, so you can run those models again
Questions freeze after the first run, because the scores describe the questions that run saw.
That rule covers those questions only. A custom format’s scoring, answer key and prompts stay
editable for the life of the benchmark, which is what makes the analyse and fix loop possible.
Prompts to try
What it will not do
- Publish, unpublish or delete a benchmark
- Add or remove phases and tasks
- Change collaborators or who may join
- Set up the AI judge, edit a run’s score by hand, or email participants
- Re-score earlier runs. They keep their old scores until each model is run again
- Create a second benchmark to fix a first. Two copies share a title and split the leaderboard
Related
Analyze the results
Find out what to fix, with proof, before you fix it.
Fixing a benchmark step by step
The walkthrough, with a real before and after.
Edit a benchmark (API)
The same edits from your own code.
Leaderboard columns (API)
Add, retitle, re-sort and hide columns programmatically.