Skip to main content
The BenchGen AI agent is not one large prompt. It is built from skills: self-contained capabilities, each covering one area of the platform, each with its own tooling and its own rules about what it may do. You never name a skill. You describe what you want, and the agent loads the one that fits.

The skills

benchgen-context has no page of its own. It is what the agent draws on when you ask what BenchGen does rather than asking it to do something.

How a skill gets picked

Your message decides. Each skill declares the things it handles, in every language the agent speaks, and the agent matches what you wrote against them: When a message could belong to two skills, the agent asks one short question instead of guessing. It answers in your language throughout, findings included.

Skills hand work to each other

This is the part that makes the agent more than a set of shortcuts. A skill that finds a problem is not the skill that fixes it, and the boundary is deliberate:
1

benchmark-analyze finds it and proves it

Read only. It can tell you the answer key is too narrow and prove it by scoring the run’s saved answers again, but it cannot change a single character of the benchmark.
2

benchmark-edit applies it

Only after you say yes, and only what was proposed. It checks the change against the real benchmark before writing anything.
3

benchmark-launch measures it again

A fresh run of the same model, so the effect of the fix is a real score and not a prediction.
The same separation runs the other way: if the analysis concludes the model is at fault rather than the benchmark, the work goes to training-launch instead. Training a model against a broken answer key would teach it the wrong answer, so deciding which of the two it is comes first.

What every skill has in common

  • It works as you. Your account, your permissions, your credits. Each chat turn gets its own short-lived credential, and the agent never sees or asks for a token.
  • It confirms before it costs. Anything that creates, changes or spends shows you what it would do and waits for your next message.
  • It talks about your benchmark, not about itself. Skills, files, endpoints and field names stay out of the conversation, which is why you will not hear a skill name in chat.
  • It has a hard ceiling. See what the agent is allowed to do for the things no skill can reach, such as deleting a benchmark or editing a score by hand.

What the agent may do

Permissions, confirmations, and the hard limits.

Working with the agent

The task guide: the full improvement loop end to end, with screenshots.
Last modified on September 24, 2026