Skip to main content

What it can do

  • Write a benchmark from a topic, a question count and a format you pick
  • Ask questions and expect one short answer, or offer options and expect one letter
  • Test whether a model picks the right tool and the right arguments
  • Expect a number, right within a tolerance you set
  • Score a free answer on the key points it mentions
  • Build the whole evaluation package for a format the standard editor cannot express
  • Upload a benchmark you already have as a zip
  • Build one from a public git repository, so the benchmark has a history
  • Drop duplicate questions and spread the correct letters across A, B, C and D
  • Rehearse a custom format against a stand-in model, all right then all wrong, before it reaches your account
  • Show every question with its answer for approval, then create the benchmark as a draft
  • Publish a draft when you ask
Chat with the BenchGen AI agent: a request for a 10-question multiple choice benchmark, the agent's preview of the questions and answers, the user's confirmation, and the agent's report that the benchmark was created as a draft with a link and three numbered next steps

From a request to a draft benchmark: the preview of every question with its answer, the confirmation, and the link

Prompts to try

What it will not do

  • Create or publish anything without an explicit request
  • Retry on its own. One create per confirmation, and a failure is reported rather than repeated
  • Change an existing benchmark, or create a second one to fix a first
  • Delete or unpublish anything
Creation is rate limited to 10 benchmarks a day and 3 a minute through the agent. The web app is not limited.

Creating a benchmark step by step

The walkthrough, with screenshots and large benchmarks.

Run a benchmark

Put a model on what you just made.

Bundle structure

What a custom format package contains.

Create from a repository (API)

Build a benchmark from a bundle in a git repository.
Last modified on September 24, 2026