Skip to main content
A BenchGen environment bundle is a .zip file. BenchGen unpacks it, reads the competition.yaml at the root, and assembles the environment from the files declared there.

Top-level layout

Files can be placed inside subdirectories — they just need to be referenced by their full relative path inside competition.yaml. Folders work in place of the inner zips. Anywhere competition.yaml names a .zip, it can instead name a folder, and BenchGen zips it for you while unpacking. This layout is equivalent to the one above and is usually easier to write and review:
Only competition.yaml itself must sit in the zip root. The two screenshots below show what a real code submission bundle (left) and dataset submission bundle (right) look like on disk:
Code submission bundle vs dataset submission bundle — top-level structure

Code submission bundle vs dataset submission bundle — top-level structure

Code submission and dataset submission contents

Code submission and dataset submission contents

Files explained

competition.yaml

The only required file at the root. It defines the environment’s metadata, tasks, phases, and leaderboard. Every other file in the bundle is referenced from here. See the YAML reference for a full field-by-field breakdown.

Reference data (reference_data.zip)

The ground-truth answers your scoring program uses to judge a model’s outputs. Only the scoring program reads this file — it is never exposed to the model being evaluated. Reference data can be anything your scoring program can parse: CSV rows, JSON objects, plain text labels, or structured prediction targets. The format is entirely up to you as long as your scoring program can read it.

Scoring program (scoring_program.zip)

The script that decides whether a model’s output is correct. BenchGen runs this after every submission. The zip must contain:
  • Your scoring script (e.g. scoring.py)
  • A metadata.yaml that specifies the command used to run it
metadata.yaml example:
BenchGen mounts the following directories when running your script: Your script must write a scores.json to /app/output/:
The keys must match the leaderboard column keys defined in competition.yaml. Any additional keys are ignored.
Your scoring program can also write a detailed_results.html to /app/output/ to display per-submission result breakdowns in the BenchGen UI.

Ingestion program (ingestion_program.zip) — optional

An ingestion program is needed when BenchGen runs the model end-to-end as part of the evaluation rather than receiving pre-generated outputs. It takes input data, calls the model, and writes predictions that the scoring program then evaluates. The zip must contain your ingestion script and a metadata.yaml file at the root of the folder:
Ingestion program folder containing metadata file

Ingestion program folder containing metadata file

The zip structure mirrors the scoring program: your script plus a metadata.yaml. metadata.yaml example:
BenchGen mounts: The argument order in metadata.yaml differs depending on your submission mode. In code submission mode the model code is $submission_program; in dataset submission mode the dataset is $submission_program and your sample code becomes $input:
Ingestion program metadata command — code submission vs dataset submission

Ingestion program metadata command — code submission vs dataset submission

How the model reaches your ingestion program

For an LLM benchmark you do not write the model, and neither does the person running it. When a run starts, BenchGen generates a model.py already configured with the endpoint, credentials, model name and sampling settings for the model under evaluation, and mounts it at /app/ingested_program/. Your ingestion program imports it:
That is the whole contract. A bundle never ships a model and never holds an API key, which is why you can write one without knowing which model it will be run against. These class names all refer to the same implementation, so a bundle written against any of them keeps working: Model, GSM8KModel, MKQAModel, TQuADModel, BilmeceModel, VisionModel, TurkishLanguageModel, TurkishGrammarModel, TwentyQuestionsModel, MMLUModel, SWEModel. Prefer Model.
Put your judgement in the scoring program, not the ingestion program. The ingestion program should ask the model and record what came back, nothing more. Scoring is the part you will want to change later, and it is the part that can be re-run without spending tokens.

Input data (input_data.zip) — optional

The test inputs handed to the ingestion program at run time. This is typically the prompt set, test features, or context documents your model needs to generate predictions. It is separate from the reference data so that the evaluation remains blind — the model sees inputs but never the answers.

Building one from zero

In order, with the reason each step exists:
  1. Decide what one item looks like, and write phase/input_data/items.json as a list of objects each with an id and whatever the model needs to see. Write the answers as phase/reference_data/items.json, each with the same id. The two files are matched by id, never by position.
  2. Write the ingestion program against the contract above: read the input items, call model.call_llm(...) for each, write the answers to /app/output/. Keep it dull.
  3. Write the scoring program: read predictions from /app/input/res/ and answers from /app/input/ref/, decide what counts as correct, write scores.json to /app/output/.
  4. Name the metrics in competition.yaml. Every leaderboard column key must be a key your scoring program actually writes into scores.json. A column whose key is never written stays empty forever, with no error anywhere. Change the two together or neither.
  5. Point competition.yaml at every file: image, terms, each page, and each task’s programs and data. Anything it names that the zip does not hold is refused at upload.
  6. Zip the contents, not the folder. See the warning at the end of this page.
  7. Dry run it with ?dry_run=1 before creating anything, and read the warnings as well as the errors.
Nothing that reads your bundle can tell you whether your scoring program is correct, only whether it is present. A scorer that marks everything correct, or that disagrees with its own reference data, passes every check here and produces a leaderboard that means nothing. Run your scoring program over your own reference data before you upload: answer every item correctly and confirm you score full marks, then answer each item with a different item’s answer and confirm the score falls. Junk input is not a useful test, because most scorers reject it before they compare anything.

Validation

When you upload a bundle, BenchGen checks:
  • competition.yaml is present at the root and parses without errors
  • All files referenced in competition.yaml exist inside the zip
  • The scoring program zip contains a metadata.yaml with a command key
  • Leaderboard column keys in competition.yaml match at least one key expected in scores.json
Validation errors are shown inline on the upload screen with the specific field or file causing the issue.

Upload from a script or an agent

Besides the upload screen, you can send a bundle straight to the API in one request with POST /api/competitions/create_from_bundle/ (multipart/form-data, the zip in the bundle field, benchmark:create scope). Add ?dry_run=1 to check a bundle without creating anything. The API reference covers the endpoint and the token scopes it needs. Before it stores anything, the API checks the archive and then the manifest. It refuses a bundle that is not a zip, has no competition.yaml in the root, has no title in that file, is larger than 512 MB, unpacks to more than 2 GB, holds more than 10000 files, compresses far beyond what real content does, or carries entries that would escape the extract folder. It then reads competition.yaml and refuses a bundle that names something the zip does not hold: no image or an image that is not there, no terms or an empty terms file, no pages or a page file that is missing or empty, no tasks, a task with no index or no scoring program, a task pointing at data that is neither in the zip nor the key of an existing dataset, or no phases. A task reference may be a dataset UUID instead of a path, and a folder reference resolves through the files under it, so neither is refused. Both the dry run and the real create also answer with warnings, which never block creation. They tell you about a leaderboard column key that no scoring program appears to write, so that column would stay empty; placeholder text left in the title or description; an empty image; and a manifest with no pages at all. Read them before you create: nothing later will mention them again. Anything past that is reported through creation_status while the bundle is unpacked, for example a bad phase date or a docker image the platform cannot run.
competition.yaml must sit in the root of the zip. Zipping the folder that contains it puts it one level down and the upload is refused. Zip the contents of the folder, not the folder itself.

Next steps

YAML Reference

All fields in competition.yaml.

Create a Custom Environment

End-to-end upload walkthrough.
Last modified on September 17, 2026