> ## Documentation Index
> Fetch the complete documentation index at: https://benchgen.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Run results by question

> Every question of a finished run: expected answer, the model's answer, right or wrong, its group, and what the scoring found missing (`missed`, such as key points) when it records that, with accuracy overall and per group. `?compare=<run id>` answers what this run fixed and broke against another run of the same benchmark instead. Readable by the person who made the run and by the benchmark's owners. 404 with code `no_item_results` when the scoring program records no per-question results. Requires `benchmark:read`.



## OpenAPI

````yaml GET /api/submissions/{id}/item_results/
openapi: 3.0.3
info:
  title: BenchGen Platform API
  version: 1.0.0
  description: >-
    One API for the whole BenchGen platform: model catalogue and serving,
    fine-tuning and benchmarks, datasets (knowledge), and billing.


    ## Authentication

    Every request uses the same credential: a platform API token sent as
    `Authorization: Bearer bgn_...`.

    Create tokens in the web app under Profile Settings > Platform API tokens
    (the secret is shown exactly once), or via `POST /api/tokens/` with an
    interactive session. Revoking a token disables it platform-wide within 60
    seconds.


    ## Scopes

    A token carries scopes chosen at creation; a request outside the token's
    scopes gets `403` with an explanatory message.


    | scope | grants |

    |---|---|

    | `models:read` | read model catalogues, job status, logs, GPU info |

    | `models:write` | deploy, train, merge, stop models and jobs |

    | `benchmark:read` | read benchmark runs and results |

    | `benchmark:run` | launch benchmark runs |

    | `benchmark:create` | create benchmarks in your account (drafts, Excel,
    bundles, specs); counts toward the creation limit |

    | `benchmark:manage` | edit and delete benchmarks you own or collaborate on
    |

    | `benchmark:publish` | publish and unpublish benchmarks you own or
    collaborate on |

    | `knowledge:read` | read your datasets and fine-tuning data |

    | `knowledge:write` | create, edit and delete datasets and fine-tuning data
    |

    | `billing:read` | read your balance and usage |

    | `agents:chat` | chat with your own agents through the API |

    | `agents:manage` | manage your agents, knowledge bases and channels |

    | `admin` | everything the account can do (staff accounts only) |


    ## For agents

    This document plus `/api/llms.txt` are the machine-readable entry points.
    Responses are JSON. Errors use conventional status codes; the body carries
    `error` or `message`. Knowledge endpoints return `[{"data": [...], "meta":
    {...}}]`.


    The complete auto-generated schema of every endpoint (including internal
    ones) lives at `/api/public-docs.json` (Swagger 2.0); this document is the
    curated, stable, supported surface.
  contact:
    url: https://benchgen.com
servers:
  - url: https://api.benchgen.com
security:
  - platformToken: []
tags:
  - name: auth
    description: Token introspection for services and integrations
  - name: tokens
    description: Manage your platform API tokens
  - name: models
    description: Model catalogue and serving
  - name: finetune
    description: Fine-tuning jobs, inference deployments, GPUs
  - name: knowledge
    description: Datasets and fine-tuning data (knowledge API)
  - name: billing
    description: Balance and usage
  - name: agents
    description: Chat with your agents (OpenAI-compatible facade)
  - name: benchmark
    description: Public benchmark (competition) listings.
paths:
  /api/submissions/{id}/item_results/:
    get:
      tags:
        - benchmark
      summary: One run, question by question
      description: >-
        Every question of a finished run: expected answer, the model's answer,
        right or wrong, its group, and what the scoring found missing (`missed`,
        such as key points) when it records that, with accuracy overall and per
        group. `?compare=<run id>` answers what this run fixed and broke against
        another run of the same benchmark instead. Readable by the person who
        made the run and by the benchmark's owners. 404 with code
        `no_item_results` when the scoring program records no per-question
        results. Requires `benchmark:read`.
      parameters:
        - name: id
          in: path
          required: true
          schema:
            type: integer
          description: The run (submission) id.
        - name: only
          in: query
          schema:
            type: string
            enum:
              - wrong
              - right
        - name: limit
          in: query
          schema:
            type: integer
            maximum: 200
        - name: offset
          in: query
          schema:
            type: integer
        - name: compare
          in: query
          schema:
            type: integer
      responses:
        '200':
          description: Per-question results
          content:
            application/json:
              schema:
                type: object
        '403':
          description: Not your run or benchmark
        '404':
          description: No such run, or no per-question results (code no_item_results)
components:
  securitySchemes:
    platformToken:
      type: http
      scheme: bearer
      bearerFormat: bgn_ opaque token
      description: >-
        Platform API token created under Profile Settings > Platform API tokens.
        Scopes are fixed at creation.

````