> ## Documentation Index
> Fetch the complete documentation index at: https://benchgen.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Benchmark an OpenClaw Agent

> Run benchmarks against a connected OpenClaw agent and read its traces in detail.

This page picks up once the plugin is installed, the **Runtime** card shows *via BenchGen relay*, and the first traces have arrived. From here you can benchmark the agent like any model and read every turn it handled in full detail.

<CardGroup cols={2}>
  <Card title="Connect an OpenClaw Agent" icon="puzzle-piece" href="/docs/agentspace/connect-openclaw-agent">
    Install the plugin and get the first traces, if you haven't yet.
  </Card>

  <Card title="Datasets from an OpenClaw Agent" icon="database" href="/docs/agentspace/datasets-openclaw-agent">
    Turn the turns you read here into dataset items.
  </Card>
</CardGroup>

## Benchmarks

Connecting the agent through the relay registers it as a model in BenchGen. From there it is benchmarked like any other model:

1. Open a benchmark from the Environments Hub (or a [custom environment](/docs/eval/create-environment)) and click **Evaluate**.

2. Under **Platform Models**, pick the agent by its name.

   <Frame caption="A benchmark's Evaluate tab: the connected OpenClaw agent picked under Platform Models">
     <img src="https://mintcdn.com/benchgen-8fc81371/oXzVpf1EVd5pntSt/images/agentspace/openclaw/08-evaluate-platform-models.png?fit=max&auto=format&n=oXzVpf1EVd5pntSt&q=85&s=afc21a70aca94924d1bc18fd20b7ea10" alt="A benchmark's Evaluate tab with a connected OpenClaw agent selected under Platform Models" width="1440" height="900" data-path="images/agentspace/openclaw/08-evaluate-platform-models.png" />
   </Frame>

3. Run. BenchGen sends each benchmark prompt to the agent through the same relay that chat uses, so the agent answers with its real system prompt, tools and model, and scores the answers as defined by the environment.

4. Watch the live log. The run opens to a **live log** view: the benchmark data loads, progress streams item by item (`Processing: 10/100 (10%)`), and the scoring output ends with the final score, for example `accuracy=90.00% correct=90/100`. The **Evaluation details** panel beside the log sums up the run.

   <Frame caption="An agent evaluation running: live logs on the left, the Evaluation details panel with the agent as the model on the right">
     <img src="https://mintcdn.com/benchgen-8fc81371/3YjgzNxekgZVURa4/images/agentspace/openclaw/08a-benchmark-run-logs.png?fit=max&auto=format&n=3YjgzNxekgZVURa4&q=85&s=a8dfd9b3e65cdabcddfb74a960048f0e" alt="An agent benchmark run with live logs and the evaluation details panel" width="1600" height="900" data-path="images/agentspace/openclaw/08a-benchmark-run-logs.png" />
   </Frame>

5. Read the run's traces. Every benchmark turn lands on the agent's **Traces** tab as a row of its own: one per task, named `benchgen`, tagged `benchgen-chat`, all sharing the session key of that run. Filter by the tag (or the run's session ID) to read, task by task, what the benchmark asked and what the agent answered.

   <Frame caption="The Traces tab filtered by tag benchgen-chat while a benchmark runs: one row per task, streaming in live">
     <img src="https://mintcdn.com/benchgen-8fc81371/3YjgzNxekgZVURa4/images/agentspace/openclaw/08b-benchmark-traces.png?fit=max&auto=format&n=3YjgzNxekgZVURa4&q=85&s=569e7e51cd9f641991d6bd2904487bca" alt="The Traces tab filtered by the benchgen-chat tag, one row per benchmark task" width="1600" height="900" data-path="images/agentspace/openclaw/08b-benchmark-traces.png" />
   </Frame>

Runs the agent took part in are listed on its **Evaluations** tab; the results page is the same as for models; see [Benchmark Results](/docs/eval/read-results).

If the Runtime card reports *Connected for chat, but BenchGen refused the registration, so this agent can't be benchmarked yet*, chat works but the model registration failed; the message carries the reason. Fix it and click **Connect now** on the card.

## Traces in detail

One trace per message the agent handles; inside it, a span per model call (with output, token usage and cost) and per tool execution (name, input, output, duration, or the error / the reason it was blocked). A trace's session key is the OpenClaw session the turn ran in, so a conversation reads as one session across turns.

The mapping from OpenClaw's diagnostics events:

| OpenClaw event | In the trace |
| - | - |
| `run.started` | trace start |
| `context.assembled` | span with the assembled context |
| `model.usage` | generation span: output, usage, cost |
| `model.call.error` | generation span closed with the error |
| `tool.execution.started` / `completed` / `error` / `blocked` | tool span: input, then output and duration, or the error / blocked reason |
| `run.completed` | trace finalized; pending spans closed |

On the **Traces** tab:

* **Time range** (last 24 h, 7 days, 30 days, all time) and filters by **trace name**, **user ID**, **session ID** and **tag**. Useful tags: `benchgen-setup` (handshakes) and `benchgen-chat` (turns that came from BenchGen: chat and benchmarks).
* Click a row to open the trace beside the list: its observations, inputs and outputs, latency, tokens and spend.
* Select rows for bulk actions: **Add to dataset**, **Add to annotation queue**, **Delete traces**.

### Open a trace

Click a trace's timestamp to open it next to the list. On the left is the tree
of observations (the agent turn, `context.assembled`, one generation per model
call, one span per tool execution) with a graph of the turn under it; on the
right, the selected observation's input, output and metadata.

<Frame caption="A benchmark turn opened: the observation tree with the turn graph, and the task's output and metadata (outcome, model, run id)">
  <img src="https://mintcdn.com/benchgen-8fc81371/3YjgzNxekgZVURa4/images/agentspace/openclaw/08c-benchmark-trace-detail.png?fit=max&auto=format&n=3YjgzNxekgZVURa4&q=85&s=7e83a3f687cb12b0f53e1bf109af0708" alt="A benchmark turn opened beside the trace list, with the observation tree, graph, output and metadata" width="1600" height="900" data-path="images/agentspace/openclaw/08c-benchmark-trace-detail.png" />
</Frame>

Select a generation in the tree to read the model call itself: the exact input
it received and its output. On plugin `0.5.0` or newer the generation also
carries the token numbers of the call (prompt tokens and context used against
the model's context window).

<Frame caption="The benchmark turn's generation selected: the exact task the model received and the answer it returned">
  <img src="https://mintcdn.com/benchgen-8fc81371/3YjgzNxekgZVURa4/images/agentspace/openclaw/08d-benchmark-trace-generation.png?fit=max&auto=format&n=3YjgzNxekgZVURa4&q=85&s=0bd62c551730943f8e56da368f6eff78" alt="A generation selected in the observation tree, showing the benchmark task as the model call's input and the answer as its output" width="1600" height="900" data-path="images/agentspace/openclaw/08d-benchmark-trace-generation.png" />
</Frame>

## Next: datasets

The turns you just read are raw material for training and annotation: selecting them on the Traces tab and choosing **Add to dataset** turns them into dataset items.

<Card title="Datasets from an OpenClaw Agent" icon="database" href="/docs/agentspace/datasets-openclaw-agent" horizontal>
  Turn real agent traces into training data.
</Card>
