Not every run has one. Whether a benchmark’s scoring program emits a finetune artifact at all, and which cases it selects, is a decision made by whoever built that environment. Runs on environments that don’t support this show “the benchmark’s scoring program did not emit one” instead, see Benchmark Results.
How a sample gets generated
The scoring program builds the dataset from the run’s own results, layering three sources:- Seed — every question the model got wrong. Each becomes one training example: the same question as input, and a short “reasoning + final answer” as the target, using the benchmark’s own correct answer, not the model’s wrong one.
- Augmentation — deterministic variants of each seed (on by default, 5x). Every wrong question is expanded into several rephrasings: the connectors reworded (“ardından” ↔ “sonra”), a different instruction prefix or suffix added, or an alternate system prompt swapped in. Numbers are never touched, so the correct answer stays valid across every variant. 40 wrong answers becomes roughly 200 training rows this way.
- Pool (optional) — a slice of a larger static dataset, when the benchmark ships one, mixed in so the model doesn’t train on failure cases alone.
selection_criterion, for example wrong_plus_augmented (seed + augmentation, the default) or all_items (every question, not just misses).
What’s inside the dataset
Unzip the artifact and you get a manifest, a short README, and the training data itself in four formats:manifest.json carries everything about how the dataset was built:
data/messages.jsonl looks like this, ready to feed straight into an SFT trainer:
suggested_hyperparams is a recommendation from whoever built the benchmark, not something Train applies automatically. Use it as a starting point when you configure the training job yourself.Getting the artifact into a dataset
Automatically
The moment a run finishes, BenchGen checks it for a finetune artifact and saves it to your datasets on its own, once per run. Most of the time there’s nothing to do.
Manually
On the run’s detail page, the Fine-tune Dataset panel below Files shows Save to FineTune Datasets whenever the auto-save hasn’t already run for that artifact.
Where it shows up
Every saved artifact becomes a dataset on the Datasets page, under the Fine-tune filter, tagged with a purple Benchmark badge. Its name is generated for you:GSM8K · Qwen2.5-7B-LoRA · 72.5% · Jun 2026
Opening it shows the metadata carried over from the run: the base model it was trained against, the number of examples, the selection criterion the scoring program used to pick them, and a View benchmark link back to the source run.
Using it in Train
Pick it like any other dataset in the Model & Dataset step of a new training job. Selecting a benchmark-sourced dataset also fills in the base model for you, since the artifact already records which model it came from, a banner confirms “Benchmark fine-tune dataset selected: training will reuse the saved BenchGen artifact and its recorded base model.”Fine-tune a Model
Start a training run with this dataset selected.
Add a Dataset
Everything else you can do on the Datasets page, including adding your own.