The skills
benchgen-context has no page of its own. It is what the agent draws on when you ask what
BenchGen does rather than asking it to do something.
How a skill gets picked
Your message decides. Each skill declares the things it handles, in every language the agent speaks, and the agent matches what you wrote against them:
When a message could belong to two skills, the agent asks one short question instead of guessing.
It answers in your language throughout, findings included.
Skills hand work to each other
This is the part that makes the agent more than a set of shortcuts. A skill that finds a problem is not the skill that fixes it, and the boundary is deliberate:1
benchmark-analyze finds it and proves it
Read only. It can tell you the answer key is too narrow and prove it by scoring the run’s
saved answers again, but it cannot change a single character of the benchmark.
2
benchmark-edit applies it
Only after you say yes, and only what was proposed. It checks the change against the real
benchmark before writing anything.
3
benchmark-launch measures it again
A fresh run of the same model, so the effect of the fix is a real score and not a prediction.
training-launch instead. Training a model against a broken
answer key would teach it the wrong answer, so deciding which of the two it is comes first.
What every skill has in common
- It works as you. Your account, your permissions, your credits. Each chat turn gets its own short-lived credential, and the agent never sees or asks for a token.
- It confirms before it costs. Anything that creates, changes or spends shows you what it would do and waits for your next message.
- It talks about your benchmark, not about itself. Skills, files, endpoints and field names stay out of the conversation, which is why you will not hear a skill name in chat.
- It has a hard ceiling. See what the agent is allowed to do for the things no skill can reach, such as deleting a benchmark or editing a score by hand.
Related
What the agent may do
Permissions, confirmations, and the hard limits.
Working with the agent
The task guide: the full improvement loop end to end, with screenshots.