> ## Documentation Index
> Fetch the complete documentation index at: https://benchgen.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Update a benchmark from a repository

> Apply the files in a public repository to a benchmark that already exists. This is the other half of keeping a benchmark in git: a change made as a commit, such as widening an answer key, becomes the benchmark's scoring on the platform. The fetch and the checks are the ones create_from_repo makes, and the install is the one replace_task_files makes, so a repository is not a weaker way in. The bundle is split into its parts by reading competition.yaml, and a part the bundle does not carry bytes for is left alone and reported in `skipped`, never guessed at. Existing runs are never re-scored: the answer lists the finished runs scored before this change so the caller can offer to run those models again. Applying is always asked for, never a consequence of a push, so an edit made in the web app cannot be overwritten by a repository nobody has looked at. Add `?dry_run=1` to see what would change, reported as `would_replace`, without changing anything. Requires the `benchmark:edit` scope; only the creator or a collaborator may do it.

The other half of keeping a benchmark in git. [Create from a
repository](/docs/api-reference/endpoint/create-benchmark-repo) builds a new
benchmark from a folder; this applies that folder to a benchmark that already
exists, so a change made as a commit, such as widening an answer key, becomes
the benchmark's scoring.

The body is the same shape: a repository, optionally a `ref` and a `path`.

```bash theme={null}
curl -X POST "https://api.benchgen.com/api/competitions/97/update_from_repo/?dry_run=1" \
  -H "Authorization: Bearer $BENCHGEN_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"repo": "https://github.com/acme/benchmarks/tree/main/benchmarks/horses"}'
```

## What it replaces, and what it leaves alone

The bundle is split into the four parts a benchmark is made of: the scoring
program, the ingestion program, the questions and the answers. Where each one
lives is read from `competition.yaml`, not assumed, so a bundle that lays its
files out differently still works.

A part the bundle carries no bytes for is **left alone and reported**, never
guessed at. That happens when `competition.yaml` does not declare it, or when
it points at a dataset already on the platform rather than a folder. The answer
names those in `skipped`.

<Note>
  Existing runs are never re-scored. Runs that finished before the change keep
  the score they got and are listed in `runs.scored_under_previous_files`, so
  you can offer to run those models again. Nothing is re-run for you, because a
  one line change to an answer key must not quietly spend a model run per entry
  on the leaderboard.
</Note>

## Nothing happens by itself

Applying is always a request you make. There is no webhook, and merging a pull
request does not change a benchmark. That is deliberate: a repository nobody has
looked at should not be able to overwrite an edit somebody made in the web app.

<Tip>
  Add `?dry_run=1` first. It answers `would_replace` with the parts that differ,
  and the runs the change would leave scored under previous files, without
  changing anything.
</Tip>

## Private repositories

The platform fetches the archive itself and holds no credentials, so only
public repositories work. A private one answers `400` with a plain explanation.

If your benchmarks live in a private repository, the BenchGen AI agent can
still apply them: it has its own working copy and sends the files up, rather
than asking the platform to fetch. See [versioning a benchmark in
git](/docs/skills/version-a-benchmark).

## Related

<CardGroup cols={2}>
  <Card title="Create from a repository" icon="code-branch" href="/docs/api-reference/endpoint/create-benchmark-repo">
    Build a new benchmark from a folder in a repository.
  </Card>

  <Card title="Replace benchmark files" icon="arrows-rotate" href="/docs/api-reference/endpoint/replace-benchmark-files">
    The same install, from an upload instead of a repository.
  </Card>
</CardGroup>


## OpenAPI

````yaml POST /api/competitions/{id}/update_from_repo/
openapi: 3.0.3
info:
  title: BenchGen Platform API
  version: 1.0.0
  description: >-
    One API for the whole BenchGen platform: model catalogue and serving,
    fine-tuning and benchmarks, datasets (knowledge), and billing.


    ## Authentication

    Every request uses the same credential: a platform API token sent as
    `Authorization: Bearer bgn_...`.

    Create tokens in the web app under Profile Settings > Platform API tokens
    (the secret is shown exactly once), or via `POST /api/tokens/` with an
    interactive session. Revoking a token disables it platform-wide within 60
    seconds.


    ## Scopes

    A token carries scopes chosen at creation; a request outside the token's
    scopes gets `403` with an explanatory message.


    | scope | grants |

    |---|---|

    | `models:read` | read model catalogues, job status, logs, GPU info |

    | `models:write` | deploy, train, merge, stop models and jobs |

    | `benchmark:read` | read benchmark runs and results |

    | `benchmark:run` | launch benchmark runs |

    | `benchmark:create` | create benchmarks in your account (drafts, Excel,
    bundles, specs); counts toward the creation limit |

    | `benchmark:manage` | edit and delete benchmarks you own or collaborate on
    |

    | `benchmark:publish` | publish and unpublish benchmarks you own or
    collaborate on |

    | `knowledge:read` | read your datasets and fine-tuning data |

    | `knowledge:write` | create, edit and delete datasets and fine-tuning data
    |

    | `billing:read` | read your balance and usage |

    | `agents:chat` | chat with your own agents through the API |

    | `agents:manage` | manage your agents, knowledge bases and channels |

    | `admin` | everything the account can do (staff accounts only) |


    ## For agents

    This document plus `/api/llms.txt` are the machine-readable entry points.
    Responses are JSON. Errors use conventional status codes; the body carries
    `error` or `message`. Knowledge endpoints return `[{"data": [...], "meta":
    {...}}]`.


    The complete auto-generated schema of every endpoint (including internal
    ones) lives at `/api/public-docs.json` (Swagger 2.0); this document is the
    curated, stable, supported surface.
  contact:
    url: https://benchgen.com
servers:
  - url: https://api.benchgen.com
security:
  - platformToken: []
tags:
  - name: auth
    description: Token introspection for services and integrations
  - name: tokens
    description: Manage your platform API tokens
  - name: models
    description: Model catalogue and serving
  - name: finetune
    description: Fine-tuning jobs, inference deployments, GPUs
  - name: knowledge
    description: Datasets and fine-tuning data (knowledge API)
  - name: billing
    description: Balance and usage
  - name: agents
    description: Chat with your agents (OpenAI-compatible facade)
  - name: benchmark
    description: Public benchmark (competition) listings.
paths:
  /api/competitions/{id}/update_from_repo/:
    post:
      tags:
        - benchmark
      summary: Update a benchmark from a repository
      description: >-
        Apply the files in a public repository to a benchmark that already
        exists. This is the other half of keeping a benchmark in git: a change
        made as a commit, such as widening an answer key, becomes the
        benchmark's scoring on the platform. The fetch and the checks are the
        ones create_from_repo makes, and the install is the one
        replace_task_files makes, so a repository is not a weaker way in. The
        bundle is split into its parts by reading competition.yaml, and a part
        the bundle does not carry bytes for is left alone and reported in
        `skipped`, never guessed at. Existing runs are never re-scored: the
        answer lists the finished runs scored before this change so the caller
        can offer to run those models again. Applying is always asked for, never
        a consequence of a push, so an edit made in the web app cannot be
        overwritten by a repository nobody has looked at. Add `?dry_run=1` to
        see what would change, reported as `would_replace`, without changing
        anything. Requires the `benchmark:edit` scope; only the creator or a
        collaborator may do it.
      parameters:
        - name: id
          in: path
          required: true
          schema:
            type: integer
          description: The benchmark id.
        - name: dry_run
          in: query
          required: false
          schema:
            type: string
            enum:
              - '1'
          description: Validate only; change nothing.
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              required:
                - repo
              properties:
                repo:
                  type: string
                  description: >-
                    The repository URL, for example
                    https://github.com/acme/benchmarks. A browse URL carrying a
                    branch and folder is accepted.
                ref:
                  type: string
                  description: >-
                    Branch, tag or commit. Defaults to the repository's default
                    branch.
                path:
                  type: string
                  description: The folder inside the repository holding competition.yaml.
                task:
                  type: string
                  description: >-
                    Which of this benchmark's tasks to change, when it has more
                    than one. The id from task_files.
                bundle_task:
                  type: integer
                  description: >-
                    Which task in the bundle to take the files from, when
                    competition.yaml declares more than one. Zero based.
      responses:
        '200':
          description: Checked (dry run) or replaced
          content:
            application/json:
              schema:
                type: object
                properties:
                  valid:
                    type: boolean
                  id:
                    type: integer
                  tasks:
                    type: array
                    items:
                      type: integer
                  files:
                    type: array
                    items:
                      type: object
                  warnings:
                    type: array
                    items:
                      type: string
                    description: >-
                      Leaderboard columns the new scoring program does not
                      appear to write.
                  replaced:
                    type: array
                    items:
                      type: string
                    description: Absent on a dry run.
                  runs:
                    type: object
                    properties:
                      finished:
                        type: integer
                      scored_under_previous_files:
                        type: array
                        description: >-
                          Finished runs scored before the latest change. They
                          keep their score; run the model again to score it with
                          the current files.
                        items:
                          type: object
                          properties:
                            id:
                              type: integer
                            model_name:
                              type: string
                            owner:
                              type: string
                            created_when:
                              type: string
                              format: date-time
        '400':
          description: Fixable input, such as a zip that is not one or a shared task
          content:
            application/json:
              schema:
                type: object
                properties:
                  valid:
                    type: boolean
                  errors:
                    type: array
                    items:
                      type: string
        '403':
          description: Not the creator or a collaborator, or the token lacks benchmark:edit
        '404':
          description: No such benchmark
components:
  securitySchemes:
    platformToken:
      type: http
      scheme: bearer
      bearerFormat: bgn_ opaque token
      description: >-
        Platform API token created under Profile Settings > Platform API tokens.
        Scopes are fixed at creation.

````