> ## Documentation Index
> Fetch the complete documentation index at: https://benchgen.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# BenchGen AI agent skills

> The BenchGen AI agent is built from skills: packaged capabilities for creating, running, analyzing and editing benchmarks and for training models.

The BenchGen AI agent is not one large prompt. It is built from **skills**: self-contained
capabilities, each covering one area of the platform, each with its own tooling and its own rules
about what it may do.

You never name a skill. You describe what you want, and the agent loads the one that fits.

## The skills

| Skill | What it covers | Changes anything |
| - | - | - |
| [`benchmark-create`](/docs/skills/create-a-benchmark) | Writing a new benchmark from a topic, or uploading one you already have | Yes, creates and publishes |
| [`benchmark-launch`](/docs/skills/run-a-benchmark) | Running a model on a benchmark and reporting the score | Yes, spends credits |
| [`benchmark-analyze`](/docs/skills/analyze-results) | Reviewing a benchmark, reading runs question by question, proposing fixes | No, read only |
| [`benchmark-edit`](/docs/skills/edit-a-benchmark) | Changing a benchmark in place: questions, scoring, leaderboard, settings | Yes, changes the benchmark |
| [`training-launch`](/docs/skills/train-a-model) | Fine-tuning, RL, distillation, and merging trained models | Yes, spends credits |
| `benchgen-context` | What BenchGen is, what it measures, who it is for, how access works | No, knowledge only |

`benchgen-context` has no page of its own. It is what the agent draws on when you ask what
BenchGen does rather than asking it to do something.

## How a skill gets picked

Your message decides. Each skill declares the things it handles, in every language the agent
speaks, and the agent matches what you wrote against them:

| You say | The skill that handles it |
| - | - |
| "create a benchmark about X", "benchmark oluştur" | `benchmark-create` |
| "run it on Claude", "how is my run?" | `benchmark-launch` |
| "why did it score so low?", "which questions failed?" | `benchmark-analyze` |
| "fix question 4", "change the scoring", "rank by accuracy" | `benchmark-edit` |
| "fine-tune it on what it got wrong" | `training-launch` |

When a message could belong to two skills, the agent asks one short question instead of guessing.
It answers in your language throughout, findings included.

## Skills hand work to each other

This is the part that makes the agent more than a set of shortcuts. A skill that finds a problem
is not the skill that fixes it, and the boundary is deliberate:

<Steps>
  <Step title="benchmark-analyze finds it and proves it">
    Read only. It can tell you the answer key is too narrow and prove it by scoring the run's
    saved answers again, but it cannot change a single character of the benchmark.
  </Step>

  <Step title="benchmark-edit applies it">
    Only after you say yes, and only what was proposed. It checks the change against the real
    benchmark before writing anything.
  </Step>

  <Step title="benchmark-launch measures it again">
    A fresh run of the same model, so the effect of the fix is a real score and not a prediction.
  </Step>
</Steps>

The same separation runs the other way: if the analysis concludes the **model** is at fault rather
than the benchmark, the work goes to `training-launch` instead. Training a model against a broken
answer key would teach it the wrong answer, so deciding which of the two it is comes first.

## What every skill has in common

* **It works as you.** Your account, your permissions, your credits. Each chat turn gets its own
  short-lived credential, and the agent never sees or asks for a token.
* **It confirms before it costs.** Anything that creates, changes or spends shows you what it
  would do and waits for your next message.
* **It talks about your benchmark, not about itself.** Skills, files, endpoints and field names
  stay out of the conversation, which is why you will not hear a skill name in chat.
* **It has a hard ceiling.** See [what the agent is allowed to do](/docs/skills/what-the-agent-may-do)
  for the things no skill can reach, such as deleting a benchmark or editing a score by hand.

## Related

<CardGroup cols={2}>
  <Card title="What the agent may do" icon="shield-halved" href="/docs/skills/what-the-agent-may-do">
    Permissions, confirmations, and the hard limits.
  </Card>

  <Card title="Working with the agent" icon="robot" href="/docs/guides/benchgen-ai-agent/working-with-the-agent">
    The task guide: the full improvement loop end to end, with screenshots.
  </Card>
</CardGroup>
