Skip to main content
Eval is BenchGen’s evaluation module: a structured, reproducible way to measure how well a model performs on a task before you ship it or fine-tune it. The core primitive is an environment, a self-contained package of test cases, a harness, and scoring rules. Pick one from the Environments Hub, point it at a model, and Eval runs and scores every case.

Run a benchmark

Point an environment at a model, whichever way that model reaches BenchGen.

On a Platform Model

Benchmark a model that is already published or deployed on BenchGen.

Evaluate an Inference Model

Benchmark a live inference endpoint and watch the run in real time.

From a HuggingFace Model

Pull a public model from the HuggingFace Hub, deploy it, and benchmark it.

From an External API

Benchmark any OpenAI-compatible endpoint, such as OpenRouter or Mistral.

Bring your own model

Get a model into BenchGen before you benchmark or serve it.

Add a Model

Register a model by uploading an archive or importing one from HuggingFace.

Deploy an Inference Model

Spin up a live, OpenAI-compatible inference endpoint for a model.

Configure and monitor a run

Environment Variables

Declare, inject, and read env vars so runs are env-var-first, no model.py to edit.

Monitor Model Usage

Track requests, tokens, latency, and spend for a deployed model.

After a run

Benchmark Results

Understand the benchmark results report and what the metrics mean.

Export Datasets → Train

Export failing cases as a labeled dataset to kick off a fine-tuning run.

Build a custom environment

If the hub doesn’t cover your task, package it yourself as a .zip bundle and upload it.

Create a Custom Environment

Build and upload your own evaluation environment to BenchGen.

Bundle Structure

The files inside a custom environment .zip bundle and what each does.

YAML Reference

Every field in competition.yaml for a custom environment bundle.
Last modified on September 3, 2026