Skip to main content
Every run, past or in progress, has a detail page. Open one from Results or Evaluations in the Eval sidebar, or from a Recents entry. While a run is still going you see live logs, covered in Evaluate an Inference Model. This page covers what you see once it’s Completed.
A completed evaluation page showing Score Breakdown, Detailed Results, and the Evaluation details and Files sidebar

A completed evaluation on GSM8K-TR: score breakdown, the environment's own detailed report, and the evaluation details and files sidebar

Score Breakdown

The Overall Score tile is followed by the individual metrics the environment’s scoring program reported. These are entirely environment-defined, not a fixed set, so two runs can show completely different tiles:
  • A math benchmark might report accuracy, exact match, answer rate, and duration.
  • A routing benchmark might instead report accuracy, correct, total, total cost usd, and avg cost usd per task.
If a scoring program doesn’t emit structured metrics, this card just says score breakdown isn’t available for that run.

Detailed Results

This embeds the HTML report the environment’s own scoring program generated, the same harness and scoring rules described in the Overview. Layout is entirely up to whoever built the environment:
  • The example above is a simple report: a few stat tiles plus a per-question table with the case ID, the Gold (expected) answer, the model’s Prediction, and a pass/fail mark.
  • Others go further, for example breaking results down by difficulty, by domain, or by routing target with their own bar charts, on top of the same accuracy-style tiles.
Click Open to view the report on its own page. It only appears once the run has a detailed-results artifact, failed runs show the failure reason and logs instead.

Evaluation details

The sidebar summarizes the run: If you own the run, a Run configuration button below the sidebar opens the sampling parameters and environment variables it used, see Environment Variables.

Files

Every artifact the run produced is downloadable from Files, each row has its own Download button so you can grab exactly the one you need:
The Files section of the run detail sidebar, highlighted, listing Submission File, Prediction Output, Scoring Output, Detailed Results, and Finetune Artifact

The Files card in the sidebar: Submission File, Prediction Output, Scoring Output, Detailed Results, and Finetune Artifact

None of these have a fixed internal structure. What ends up inside Prediction Output, Scoring Output, and Detailed Results is entirely up to the benchmark’s own scoring_program.zip (and ingestion_program.zip, if it has one), see Bundle Structure. The examples below show what BenchGen’s own built-in templates write, a custom environment can look completely different.

Prediction Output example

BenchGen’s Question-Answer template’s ingestion program writes one predictions.json per run:

Scoring Output example

The Code (SWE-style) template’s scoring program reads that predictions.json and writes its own scores.json, with whatever metric keys that benchmark defines:
These scores.json keys are exactly what populates the Score Breakdown tiles for that run, matched against the leaderboard columns declared in the environment’s competition.yaml.

Finetune Artifact example

When a scoring program also writes a finetune_artifact.zip, BenchGen extracts it and lists it here. Inside, it’s a small manifest plus the training rows:
Below Files, a Fine-tune Dataset panel either lets you save this artifact as a training dataset, or tells you the scoring program didn’t emit one for this run. Turning results into training data is covered in Export Datasets → Train.
Last modified on September 3, 2026