Skip to main content
The BenchGen platform exposes one API surface: model catalogue and serving, fine-tuning, benchmarks, datasets (knowledge) and billing. Every authenticated endpoint accepts the same credential, a platform API token.

Base URL

One host serves the whole surface. Knowledge (dataset) endpoints live under https://api.benchgen.com/knowledge/, everything else is served at the root.

Authentication

Create a token in the web app under Profile Settings > Platform API tokens. The secret is shown exactly once, only its hash is stored. Send it on every request:
Creating and revoking tokens always requires an interactive sign-in: the web app, or an interactive token from POST /api/api-token-auth/ (username or email plus password). A platform token cannot mint or revoke tokens, so a leaked token can never widen its own reach. Revoking a token disables it platform-wide within 60 seconds.

Scopes

A token carries scopes chosen at creation. A request outside the token’s scopes returns 403 with the missing scope named in the message.

Public endpoints

Two endpoints need no authentication at all: listing published benchmarks and listing public models. Anonymous responses are deliberately narrower than what a token gets, credential fields are always stripped.

Create and run a benchmark

A typical script or agent flow, with a token that has benchmark:create, benchmark:publish, benchmark:run and benchmark:read:
  1. Check the spec. POST /api/competitions/create_from_spec/?dry_run=1 with a title and 5-2000 question-answer or multiple-choice items. Fix any errors it lists; nothing is created.
  2. Create. Send the same body without dry_run. Poll GET /api/competitions/{status_id}/creation_status/ until status is Finished; competition then carries the benchmark id.
  3. Publish. POST /api/competitions/{id}/toggle_publish/ with {"published": true}.
  4. Run a model. POST /api/competitions/{id}/register/, then POST /models/api/queue_benchmark_run. The answer returns right away with submission_id and run_status_path.
  5. Follow the run. GET /api/submissions/{id}/run_status/ answers state, done, score and run_path. Wait check_again_seconds between calls and stop when done is true. Lost the id? GET /api/submissions/my_runs/?competition={id} lists your latest runs.
For API tokens and agents, benchmark creation is limited per user (by default 10 per rolling 24 hours and 3 per minute). Over the limit the API answers 429 with a Retry-After header; wait that long before trying again. Dry runs do not count.

Upload a bundle instead

A spec is the quick path, and it only reaches question-answer and multiple-choice benchmarks. A bundle is the full format the platform builds underneath, so it is how you create anything the spec cannot express:
  • several phases, each with its own dates and submission limits
  • several tasks per phase
  • your own scoring program, and therefore your own metrics
  • leaderboard columns you name, sort and aggregate yourself
  • image and code benchmarks
  • a custom Docker image to run submissions in
Bundle structure covers the files and YAML reference covers every competition.yaml field. If you already have one, POST /api/competitions/create_from_bundle/ takes it in one request: multipart/form-data with the zip in the bundle field. It accepts the same ?dry_run=1, answers the same 400 with an errors list, and returns the same status_id to poll. From there the flow is identical.
The YAML reference ends with a complete working competition.yaml. Starting from that is quicker than writing one from scratch.
competition.yaml has to be in the root of the zip. Zipping the folder that contains it puts it one level down and the upload is refused. Zip the contents of the folder, not the folder itself.
A bundle is also refused if it is not a zip, has no title in its manifest, is larger than 512 MB, unpacks to more than 2 GB, holds more than 10000 files, compresses far beyond what real content does, or carries entries that would escape the extract folder. Everything else about the bundle is checked while it is unpacked and reported through creation_status. The web app’s three-step presigned upload still exists and is what a browser should use; this endpoint is the one-request path for scripts and agents, and it charges the same creation limits.

Edit a benchmark

GET /api/competitions/{id}/spec/ gives a benchmark back in the same item shape create_from_spec accepts, so you can read it, change what you need and send it back with POST /api/competitions/{id}/edit_spec/ (benchmark:edit). Send any of title, description, terms, pages, contact_email, organization_name, reward, report and items; leave out what stays the same. Everything except items is documentation. It changes in place and stays editable however much has been run against the benchmark. pages are the benchmark’s markdown tabs, sent in display order as [{"title": "...", "content": "..."}], at most 20 of them, and a benchmark cannot be left with none.
pages and items each replace their whole list, they are not merged into what is there. Send back every page or question you want the benchmark to keep. Read the current ones from spec/ first.
Questions can only be replaced while nothing has been run against the benchmark. A finished run’s score describes the questions that run saw, and a run in flight would read a dataset that changed underneath it, so both answer 409 with a runs count. Create a new benchmark with the corrected questions instead; retrying will not help. The title and description stay editable after a run, and editing never changes whether the benchmark is published. Older benchmarks, and ones made from an uploaded bundle, answer spec/ with items_readable: false and an empty items list: the platform cannot reconstruct their questions from a spec. Their text and pages still change through edit_spec/, and a bundle benchmark’s scoring and data change through its files, below.

Replace a bundle benchmark’s files

A benchmark built from a bundle keeps its judgement in its own scoring program and its answers in its reference data. You change those in place, with no new benchmark:
  1. GET /api/competitions/{id}/task_files/ (benchmark:read) lists each task and its scoring_program, ingestion_program, reference_data and input_data, each with a download_url that stays valid for an hour. Download the zip from that link without an Authorization header: the link carries its own.
  2. Change what you need and zip each changed part again, its contents at the root. A program zip needs its metadata.yaml at the root.
  3. POST /api/competitions/{id}/replace_task_files/ (benchmark:edit) as multipart/form-data, one zip per part under its field name, plus task when the benchmark has more than one. Add ?dry_run=1 to check first.
Unlike a question-answer benchmark’s questions, these files stay replaceable however much has been run. Replacing never touches a run: a run that finished before the change keeps its score, and the answer lists it under runs.scored_under_previous_files with its model, so you can run that model again for a score under the new files. Nothing is re-run for you, because a run spends model calls. warnings names any leaderboard column the new scoring program does not appear to write.
A task that another benchmark also uses is refused, because changing it would change that benchmark too.

Add a scoring type

A finished run’s scores.json reaches the leaderboard key by key, and a key with no column of the same name is dropped. So when a scoring program gains a metric, give the leaderboard a column for it with POST /api/competitions/{id}/leaderboard_columns/ (benchmark:edit):
A new key is added at the end. Runs scored before it show it blank; later runs fill it in. The same call changes an existing column’s title, sorting or hidden. A column is never deleted and its key never changes, because the key links it to the scores already stored; hide it instead. The column the board ranks by cannot be hidden. GET the same path to see the columns first. To rank the board by another column, send {"ranks_by": "<key>"} on its own. Every existing result moves at once, so the answer shows the top of the board now and after; send it with ?dry_run=1 first.

Settings and phases

GET /api/competitions/{id}/settings/ (benchmark:read) shows a benchmark’s switches, docker image and phases. POST the same path (benchmark:edit) changes the switches (such as enable_detailed_results) or the docker_image, or a logo image sent on its own as multipart/form-data. POST /api/competitions/{id}/phase/ changes one phase’s name, dates, submission limits, time limit and output hiding; phases must stay one after another. The AI judge settings, collaborators and the whitelist stay in the web app.

Analyse a benchmark

GET /api/submissions/{id}/item_results/ (benchmark:read) shows one run question by question: the expected answer, the model’s answer, whether it was right, and its group, with accuracy overall and per group. ?only=wrong narrows it to the misses; ?compare=<run id> answers what this run fixed and broke against another run of the same benchmark. GET /api/competitions/{id}/item_stats/ shows every question across the latest runs. A question every run misses usually has a wrong or too narrow answer key; a question every run gets right cannot tell models apart. Per-question results come from the scoring program. Question-answer and multiple-choice benchmarks record them; a custom-format bundle does when its scoring program writes finetune_artifact.zip, which the benchmark-create scaffold does. Otherwise these answer 404 with code no_item_results.

For agents and tooling

The API host serves machine-readable descriptions of itself:
  • GET /openapi.json (also /api/openapi.json), the curated OpenAPI 3 document (same surface as these pages)
  • GET /llms.txt (also /api/llms.txt), a plain-text index for LLM agents
  • GET /, a redirect to the interactive Swagger UI
Errors follow convention: 401 means the token is missing, revoked or expired (do not retry it), 403 names the missing scope, 400 explains a malformed request.
Last modified on September 21, 2026