Skip to main content
A benchmark run evaluates a model against every test case in an environment and produces a scored results report. The flow is the same whatever the model is; the only thing that changes is where the model comes from, which you pick on the Evaluate tab.

Pick your model source

Open a benchmark, click Evaluate, and choose the tab that matches your model:

The shared flow

Every source follows the same three steps once a model is selected:
1

Select a model

On the Evaluate tab, pick the source and choose your model. The Run Evaluation button activates once a model is selected.
2

Set run values (optional)

Expand Advanced: model environment & parameters. Model connection values are auto-filled from the selected model; you fill only the fields the benchmark owner exposed, plus sampling parameters. See Environment variables.
3

Run and review

Click Run Evaluation. BenchGen streams live logs and computes the score. The completed run shows a Score Breakdown and a per-question Detailed Results table. See Read results.

Tips

  • Run the same benchmark against several models to compare them on the leaderboard.
  • Start with a small benchmark (10 to 50 cases) to validate your setup before scaling up.
  • Re-run the same configuration after each training iteration to track progress.

Next steps

Last modified on July 14, 2026