Skip to main content
POST
Update a benchmark from a repository
The other half of keeping a benchmark in git. Create from a repository builds a new benchmark from a folder; this applies that folder to a benchmark that already exists, so a change made as a commit, such as widening an answer key, becomes the benchmark’s scoring. The body is the same shape: a repository, optionally a ref and a path.

What it replaces, and what it leaves alone

The bundle is split into the four parts a benchmark is made of: the scoring program, the ingestion program, the questions and the answers. Where each one lives is read from competition.yaml, not assumed, so a bundle that lays its files out differently still works. A part the bundle carries no bytes for is left alone and reported, never guessed at. That happens when competition.yaml does not declare it, or when it points at a dataset already on the platform rather than a folder. The answer names those in skipped.
Existing runs are never re-scored. Runs that finished before the change keep the score they got and are listed in runs.scored_under_previous_files, so you can offer to run those models again. Nothing is re-run for you, because a one line change to an answer key must not quietly spend a model run per entry on the leaderboard.

Nothing happens by itself

Applying is always a request you make. There is no webhook, and merging a pull request does not change a benchmark. That is deliberate: a repository nobody has looked at should not be able to overwrite an edit somebody made in the web app.
Add ?dry_run=1 first. It answers would_replace with the parts that differ, and the runs the change would leave scored under previous files, without changing anything.

Private repositories

The platform fetches the archive itself and holds no credentials, so only public repositories work. A private one answers 400 with a plain explanation. If your benchmarks live in a private repository, the BenchGen AI agent can still apply them: it has its own working copy and sends the files up, rather than asking the platform to fetch. See versioning a benchmark in git.

Create from a repository

Build a new benchmark from a folder in a repository.

Replace benchmark files

The same install, from an upload instead of a repository.

Authorizations

Authorization
string
header
required

Platform API token created under Profile Settings > Platform API tokens. Scopes are fixed at creation.

Path Parameters

id
integer
required

The benchmark id.

Query Parameters

dry_run
enum<string>

Validate only; change nothing.

Available options:
1

Body

application/json
repo
string
required

The repository URL, for example https://github.com/acme/benchmarks. A browse URL carrying a branch and folder is accepted.

ref
string

Branch, tag or commit. Defaults to the repository's default branch.

path
string

The folder inside the repository holding competition.yaml.

task
string

Which of this benchmark's tasks to change, when it has more than one. The id from task_files.

bundle_task
integer

Which task in the bundle to take the files from, when competition.yaml declares more than one. Zero based.

Response

Checked (dry run) or replaced

valid
boolean
id
integer
tasks
integer[]
files
object[]
warnings
string[]

Leaderboard columns the new scoring program does not appear to write.

replaced
string[]

Absent on a dry run.

runs
object
Last modified on September 29, 2026