What it can do
- Record a benchmark’s files in a git repository, one commit per change
- Write the commit message from the change itself, so the log reads as a history
- Record nothing when the repository already matches, rather than making an empty commit
- List every version of a benchmark, newest first
- Put an earlier version back, after a check that says what would change
- Send only the parts that actually differ, so no run is marked stale for nothing
- Keep a benchmark’s versions through a retitle, because a version is matched by its number
Why it exists
A benchmark on the platform has a timestamp, not a version. The platform can tell you a run is older than the current files. It cannot tell you what changed between them, or give the old files back. Recorded in git, a benchmark gets both. Widening an answer key becomes a diff somebody can read, and a score has a commit to point at instead of a date.Prompts to try
Where a benchmark is filed
What it will not do
- Create a benchmark, launch a run or start training
- Put a version back without a check first, which you see before anything changes
- Ask you for a repository address, a token or a key. Those are deployment settings, and a token pasted into a chat is a leaked token
- Work at all if the deployment has no repository configured. It says so plainly
The history is only as good as the exports. A change that is applied but never
recorded is a version that does not exist, which is why the agent records one
straight after applying a change, using the same words that described it.
Related
Edit a benchmark
Make the change that is then recorded.
Analyze the results
Decide what to change, with proof, before changing it.
Update from a repository (API)
Apply a repository’s files from your own code.
Create from a repository (API)
Build a new benchmark from a repository.