Skip to main content
When a benchmark’s scoring program packages a finetune artifact for a run, that’s a ready-to-train dataset built from the run’s own cases. Today that’s built into BenchGen’s Turkish benchmark suite (GSM8K-TR, Bilmece, MKQA, TQuAD, and SSB); other environments can add the same capability, see Bundle Structure if you’re the one authoring one.
Not every run has one. Whether a benchmark’s scoring program emits a finetune artifact at all, and which cases it selects, is a decision made by whoever built that environment. Runs on environments that don’t support this show “the benchmark’s scoring program did not emit one” instead, see Benchmark Results.

How a sample gets generated

The scoring program builds the dataset from the run’s own results, layering three sources:
  1. Seed — every question the model got wrong. Each becomes one training example: the same question as input, and a short “reasoning + final answer” as the target, using the benchmark’s own correct answer, not the model’s wrong one.
  2. Augmentation — deterministic variants of each seed (on by default, 5x). Every wrong question is expanded into several rephrasings: the connectors reworded (“ardından” ↔ “sonra”), a different instruction prefix or suffix added, or an alternate system prompt swapped in. Numbers are never touched, so the correct answer stays valid across every variant. 40 wrong answers becomes roughly 200 training rows this way.
  3. Pool (optional) — a slice of a larger static dataset, when the benchmark ships one, mixed in so the model doesn’t train on failure cases alone.
Which of these are combined is recorded as the manifest’s selection_criterion, for example wrong_plus_augmented (seed + augmentation, the default) or all_items (every question, not just misses).

What’s inside the dataset

Unzip the artifact and you get a manifest, a short README, and the training data itself in four formats:
manifest.json carries everything about how the dataset was built:
And each row in data/messages.jsonl looks like this, ready to feed straight into an SFT trainer:
suggested_hyperparams is a recommendation from whoever built the benchmark, not something Train applies automatically. Use it as a starting point when you configure the training job yourself.

Getting the artifact into a dataset

Automatically

The moment a run finishes, BenchGen checks it for a finetune artifact and saves it to your datasets on its own, once per run. Most of the time there’s nothing to do.

Manually

On the run’s detail page, the Fine-tune Dataset panel below Files shows Save to FineTune Datasets whenever the auto-save hasn’t already run for that artifact.
Either way, the panel then shows Open Dataset so you can jump straight to it.

Where it shows up

Every saved artifact becomes a dataset on the Datasets page, under the Fine-tune filter, tagged with a purple Benchmark badge. Its name is generated for you:
For example: GSM8K · Qwen2.5-7B-LoRA · 72.5% · Jun 2026 Opening it shows the metadata carried over from the run: the base model it was trained against, the number of examples, the selection criterion the scoring program used to pick them, and a View benchmark link back to the source run.

Using it in Train

Pick it like any other dataset in the Model & Dataset step of a new training job. Selecting a benchmark-sourced dataset also fills in the base model for you, since the artifact already records which model it came from, a banner confirms “Benchmark fine-tune dataset selected: training will reuse the saved BenchGen artifact and its recorded base model.”

Fine-tune a Model

Start a training run with this dataset selected.

Add a Dataset

Everything else you can do on the Datasets page, including adding your own.
Last modified on September 3, 2026