Skip to main content
This page picks up once the plugin is installed, the Runtime card shows via BenchGen relay, and the first traces have arrived. From here you can benchmark the agent like any model and read every turn it handled in full detail.

Connect an OpenClaw Agent

Install the plugin and get the first traces, if you haven’t yet.

Datasets from an OpenClaw Agent

Turn the turns you read here into dataset items.

Benchmarks

Connecting the agent through the relay registers it as a model in BenchGen. From there it is benchmarked like any other model:
  1. Open a benchmark from the Environments Hub (or a custom environment) and click Evaluate.
  2. Under Platform Models, pick the agent by its name.
    A benchmark's Evaluate tab with a connected OpenClaw agent selected under Platform Models

    A benchmark's Evaluate tab: the connected OpenClaw agent picked under Platform Models

  3. Run. BenchGen sends each benchmark prompt to the agent through the same relay that chat uses, so the agent answers with its real system prompt, tools and model, and scores the answers as defined by the environment.
  4. Watch the live log. The run opens to a live log view: the benchmark data loads, progress streams item by item (Processing: 10/100 (10%)), and the scoring output ends with the final score, for example accuracy=90.00% correct=90/100. The Evaluation details panel beside the log sums up the run.
    An agent benchmark run with live logs and the evaluation details panel

    An agent evaluation running: live logs on the left, the Evaluation details panel with the agent as the model on the right

  5. Read the run’s traces. Every benchmark turn lands on the agent’s Traces tab as a row of its own: one per task, named benchgen, tagged benchgen-chat, all sharing the session key of that run. Filter by the tag (or the run’s session ID) to read, task by task, what the benchmark asked and what the agent answered.
    The Traces tab filtered by the benchgen-chat tag, one row per benchmark task

    The Traces tab filtered by tag benchgen-chat while a benchmark runs: one row per task, streaming in live

Runs the agent took part in are listed on its Evaluations tab; the results page is the same as for models; see Benchmark Results. If the Runtime card reports Connected for chat, but BenchGen refused the registration, so this agent can’t be benchmarked yet, chat works but the model registration failed; the message carries the reason. Fix it and click Connect now on the card.

Traces in detail

One trace per message the agent handles; inside it, a span per model call (with output, token usage and cost) and per tool execution (name, input, output, duration, or the error / the reason it was blocked). A trace’s session key is the OpenClaw session the turn ran in, so a conversation reads as one session across turns. The mapping from OpenClaw’s diagnostics events: On the Traces tab:
  • Time range (last 24 h, 7 days, 30 days, all time) and filters by trace name, user ID, session ID and tag. Useful tags: benchgen-setup (handshakes) and benchgen-chat (turns that came from BenchGen: chat and benchmarks).
  • Click a row to open the trace beside the list: its observations, inputs and outputs, latency, tokens and spend.
  • Select rows for bulk actions: Add to dataset, Add to annotation queue, Delete traces.

Open a trace

Click a trace’s timestamp to open it next to the list. On the left is the tree of observations (the agent turn, context.assembled, one generation per model call, one span per tool execution) with a graph of the turn under it; on the right, the selected observation’s input, output and metadata.
A benchmark turn opened beside the trace list, with the observation tree, graph, output and metadata

A benchmark turn opened: the observation tree with the turn graph, and the task's output and metadata (outcome, model, run id)

Select a generation in the tree to read the model call itself: the exact input it received and its output. On plugin 0.5.0 or newer the generation also carries the token numbers of the call (prompt tokens and context used against the model’s context window).
A generation selected in the observation tree, showing the benchmark task as the model call's input and the answer as its output

The benchmark turn's generation selected: the exact task the model received and the answer it returned

Next: datasets

The turns you just read are raw material for training and annotation: selecting them on the Traces tab and choosing Add to dataset turns them into dataset items.

Datasets from an OpenClaw Agent

Turn real agent traces into training data.
Last modified on September 4, 2026