Run a benchmark
Point an environment at a model, whichever way that model reaches BenchGen.On a Platform Model
Benchmark a model that is already published or deployed on BenchGen.
Evaluate an Inference Model
Benchmark a live inference endpoint and watch the run in real time.
From a HuggingFace Model
Pull a public model from the HuggingFace Hub, deploy it, and benchmark it.
From an External API
Benchmark any OpenAI-compatible endpoint, such as OpenRouter or Mistral.
Bring your own model
Get a model into BenchGen before you benchmark or serve it.Add a Model
Register a model by uploading an archive or importing one from HuggingFace.
Deploy an Inference Model
Spin up a live, OpenAI-compatible inference endpoint for a model.
Configure and monitor a run
Environment Variables
Declare, inject, and read env vars so runs are env-var-first, no model.py to edit.
Monitor Model Usage
Track requests, tokens, latency, and spend for a deployed model.
After a run
Benchmark Results
Understand the benchmark results report and what the metrics mean.
Export Datasets → Train
Export failing cases as a labeled dataset to kick off a fine-tuning run.
Build a custom environment
If the hub doesn’t cover your task, package it yourself as a.zip bundle and upload it.
Create a Custom Environment
Build and upload your own evaluation environment to BenchGen.
Bundle Structure
The files inside a custom environment
.zip bundle and what each does.YAML Reference
Every field in
competition.yaml for a custom environment bundle.