Base URL
https://api.benchgen.com/knowledge/, everything else is served at the root.
Authentication
Create a token in the web app under Profile Settings > Platform API tokens. The secret is shown exactly once, only its hash is stored. Send it on every request:Creating and revoking tokens always requires an interactive sign-in: the web
app, or an interactive token from
POST /api/api-token-auth/ (username or
email plus password). A platform token cannot mint or revoke tokens, so a
leaked token can never widen its own reach. Revoking a token disables it
platform-wide within 60 seconds.Scopes
A token carries scopes chosen at creation. A request outside the token’s scopes returns403 with the missing scope named in the message.
Public endpoints
Two endpoints need no authentication at all: listing published benchmarks and listing public models. Anonymous responses are deliberately narrower than what a token gets, credential fields are always stripped.Create and run a benchmark
A typical script or agent flow, with a token that hasbenchmark:create,
benchmark:publish, benchmark:run and benchmark:read:
- Check the spec.
POST /api/competitions/create_from_spec/?dry_run=1with a title and 5-2000 question-answer or multiple-choice items. Fix anyerrorsit lists; nothing is created. - Create. Send the same body without
dry_run. PollGET /api/competitions/{status_id}/creation_status/untilstatusisFinished;competitionthen carries the benchmark id. - Publish.
POST /api/competitions/{id}/toggle_publish/with{"published": true}. - Run a model.
POST /api/competitions/{id}/register/, thenPOST /models/api/queue_benchmark_run. The answer returns right away withsubmission_idandrun_status_path. - Follow the run.
GET /api/submissions/{id}/run_status/answersstate,done,scoreandrun_path. Waitcheck_again_secondsbetween calls and stop whendoneis true. Lost the id?GET /api/submissions/my_runs/?competition={id}lists your latest runs.
For API tokens and agents, benchmark creation is limited per user (by default
10 per rolling 24 hours and 3 per minute). Over the limit the API answers
429
with a Retry-After header; wait that long before trying again. Dry runs do
not count.Upload a bundle instead
A spec is the quick path, and it only reaches question-answer and multiple-choice benchmarks. A bundle is the full format the platform builds underneath, so it is how you create anything the spec cannot express:- several phases, each with its own dates and submission limits
- several tasks per phase
- your own scoring program, and therefore your own metrics
- leaderboard columns you name, sort and aggregate yourself
- image and code benchmarks
- a custom Docker image to run submissions in
competition.yaml field.
If you already have one, POST /api/competitions/create_from_bundle/ takes it
in one request: multipart/form-data with the zip in the bundle field. It
accepts the same ?dry_run=1, answers the same 400 with an errors list,
and returns the same status_id to poll. From there the flow is identical.
A bundle is also refused if it is not a zip, has no title in its manifest, is
larger than 512 MB, unpacks to more than 2 GB, holds more than 10000 files,
compresses far beyond what real content does, or carries entries that would
escape the extract folder. Everything else about the bundle is checked while
it is unpacked and reported through creation_status.
The web app’s three-step presigned upload still exists and is what a browser
should use; this endpoint is the one-request path for scripts and agents, and
it charges the same creation limits.
Edit a benchmark
GET /api/competitions/{id}/spec/ gives a benchmark back in the same item
shape create_from_spec accepts, so you can read it, change what you need and
send it back with POST /api/competitions/{id}/edit_spec/ (benchmark:edit).
Send any of title, description, terms, pages, contact_email,
organization_name, reward, report and items; leave out what stays the
same.
Everything except items is documentation. It changes in place and stays
editable however much has been run against the benchmark. pages are the
benchmark’s markdown tabs, sent in display order as
[{"title": "...", "content": "..."}], at most 20 of them, and a benchmark
cannot be left with none.
Questions can only be replaced while nothing has been run against the
benchmark. A finished run’s score describes the questions that run saw, and a
run in flight would read a dataset that changed underneath it, so both answer
409 with a runs count. Create a new benchmark with the corrected questions
instead; retrying will not help. The title and description stay editable after
a run, and editing never changes whether the benchmark is published.
Older benchmarks, and ones made from an uploaded bundle, answer spec/ with
items_readable: false and an empty items list: the platform cannot
reconstruct their questions from a spec. Their text and pages still change
through edit_spec/, and a bundle benchmark’s scoring and data change through
its files, below.
Replace a bundle benchmark’s files
A benchmark built from a bundle keeps its judgement in its own scoring program and its answers in its reference data. You change those in place, with no new benchmark:GET /api/competitions/{id}/task_files/(benchmark:read) lists each task and itsscoring_program,ingestion_program,reference_dataandinput_data, each with adownload_urlthat stays valid for an hour. Download the zip from that link without anAuthorizationheader: the link carries its own.- Change what you need and zip each changed part again, its contents at the
root. A program zip needs its
metadata.yamlat the root. POST /api/competitions/{id}/replace_task_files/(benchmark:edit) asmultipart/form-data, one zip per part under its field name, plustaskwhen the benchmark has more than one. Add?dry_run=1to check first.
runs.scored_under_previous_files with its model, so you can run that model
again for a score under the new files. Nothing is re-run for you, because a
run spends model calls. warnings names any leaderboard column the new scoring
program does not appear to write.
A task that another benchmark also uses is refused, because changing it would
change that benchmark too.
Add a scoring type
A finished run’sscores.json reaches the leaderboard key by key, and a key
with no column of the same name is dropped. So when a scoring program gains a
metric, give the leaderboard a column for it with
POST /api/competitions/{id}/leaderboard_columns/ (benchmark:edit):
title, sorting or
hidden. A column is never deleted and its key never changes, because the key
links it to the scores already stored; hide it instead. The column the board
ranks by cannot be hidden. GET the same path to see the columns first.
To rank the board by another column, send {"ranks_by": "<key>"} on its own.
Every existing result moves at once, so the answer shows the top of the board
now and after; send it with ?dry_run=1 first.
Settings and phases
GET /api/competitions/{id}/settings/ (benchmark:read) shows a benchmark’s
switches, docker image and phases. POST the same path (benchmark:edit)
changes the switches (such as enable_detailed_results) or the
docker_image, or a logo image sent on its own as multipart/form-data.
POST /api/competitions/{id}/phase/ changes one phase’s name, dates,
submission limits, time limit and output hiding; phases must stay one after
another. The AI judge settings, collaborators and the whitelist stay in the
web app.
Analyse a benchmark
GET /api/submissions/{id}/item_results/ (benchmark:read) shows one run
question by question: the expected answer, the model’s answer, whether it was
right, and its group, with accuracy overall and per group. ?only=wrong
narrows it to the misses; ?compare=<run id> answers what this run fixed and
broke against another run of the same benchmark.
GET /api/competitions/{id}/item_stats/ shows every question across the
latest runs. A question every run misses usually has a wrong or too narrow
answer key; a question every run gets right cannot tell models apart.
Per-question results come from the scoring program. Question-answer and
multiple-choice benchmarks record them; a custom-format bundle does when its
scoring program writes finetune_artifact.zip, which the benchmark-create
scaffold does. Otherwise these answer 404 with code no_item_results.
For agents and tooling
The API host serves machine-readable descriptions of itself:GET /openapi.json(also/api/openapi.json), the curated OpenAPI 3 document (same surface as these pages)GET /llms.txt(also/api/llms.txt), a plain-text index for LLM agentsGET /, a redirect to the interactive Swagger UI
401 means the token is missing, revoked or expired
(do not retry it), 403 names the missing scope, 400 explains a malformed
request.