Benchgen

BenchGen Agent Check

1 phaseActive

30 chat requests a platform user makes, answered by the BenchGen AI agent about its own account. Each reply is scored on what it must contain. Portable: no fixed ids, so it runs on any environment.

Overview

BenchGen Agent Check

Thirty chat requests a user of the BenchGen platform makes, answered by the BenchGen AI agent acting as one account. Every reply is scored on what it must contain: a real score, a real run link (an /evaluations/ path), the dataset rows the user asked to see, a diagnosis, or the confirmation question before a paid action.

The tasks refer to the account's own latest objects ("my latest finished run", "my most recent benchmark"), so the benchmark runs on any environment where that account owns at least one published benchmark with a finished run. This is a check for harness changes (the agent's rules and skills) and for model changes: run it before and after, compare.