> ## Documentation Index
> Fetch the complete documentation index at: https://benchgen.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Create a benchmark

> One request, no files to build: send a title, a description and 5-2000 items. The platform validates them, builds the benchmark with its question-answer or multiple-choice importer (which generates the ingestion, scoring and LLM-judge programs) and queues creation. Items with `choices` make a multiple-choice benchmark; otherwise it is question-answer. The benchmark is created unpublished. Requires the `benchmark:create` scope. For API tokens and agents, creation is limited per user across every creation path (default 10 per rolling 24 hours and 3 per minute); a dry run does not count, and the web app and platform admins are not limited.



## OpenAPI

````yaml POST /api/competitions/create_from_spec/
openapi: 3.0.3
info:
  title: BenchGen Platform API
  version: 1.0.0
  description: >-
    One API for the whole BenchGen platform: model catalogue and serving,
    fine-tuning and benchmarks, datasets (knowledge), and billing.


    ## Authentication

    Every request uses the same credential: a platform API token sent as
    `Authorization: Bearer bgn_...`.

    Create tokens in the web app under Profile Settings > Platform API tokens
    (the secret is shown exactly once), or via `POST /api/tokens/` with an
    interactive session. Revoking a token disables it platform-wide within 60
    seconds.


    ## Scopes

    A token carries scopes chosen at creation; a request outside the token's
    scopes gets `403` with an explanatory message.


    | scope | grants |

    |---|---|

    | `models:read` | read model catalogues, job status, logs, GPU info |

    | `models:write` | deploy, train, merge, stop models and jobs |

    | `benchmark:read` | read benchmark runs and results |

    | `benchmark:run` | launch benchmark runs |

    | `benchmark:create` | create benchmarks in your account (drafts, Excel,
    bundles, specs); counts toward the creation limit |

    | `benchmark:manage` | edit and delete benchmarks you own or collaborate on
    |

    | `benchmark:publish` | publish and unpublish benchmarks you own or
    collaborate on |

    | `knowledge:read` | read your datasets and fine-tuning data |

    | `knowledge:write` | create, edit and delete datasets and fine-tuning data
    |

    | `billing:read` | read your balance and usage |

    | `agents:chat` | chat with your own agents through the API |

    | `agents:manage` | manage your agents, knowledge bases and channels |

    | `admin` | everything the account can do (staff accounts only) |


    ## For agents

    This document plus `/api/llms.txt` are the machine-readable entry points.
    Responses are JSON. Errors use conventional status codes; the body carries
    `error` or `message`. Knowledge endpoints return `[{"data": [...], "meta":
    {...}}]`.


    The complete auto-generated schema of every endpoint (including internal
    ones) lives at `/api/public-docs.json` (Swagger 2.0); this document is the
    curated, stable, supported surface.
  contact:
    url: https://benchgen.com
servers:
  - url: https://api.benchgen.com
security:
  - platformToken: []
tags:
  - name: auth
    description: Token introspection for services and integrations
  - name: tokens
    description: Manage your platform API tokens
  - name: models
    description: Model catalogue and serving
  - name: finetune
    description: Fine-tuning jobs, inference deployments, GPUs
  - name: knowledge
    description: Datasets and fine-tuning data (knowledge API)
  - name: billing
    description: Balance and usage
  - name: agents
    description: Chat with your agents (OpenAI-compatible facade)
  - name: benchmark
    description: Public benchmark (competition) listings.
paths:
  /api/competitions/create_from_spec/:
    post:
      tags:
        - benchmark
      summary: Create a benchmark from a JSON spec
      description: >-
        One request, no files to build: send a title, a description and 5-2000
        items. The platform validates them, builds the benchmark with its
        question-answer or multiple-choice importer (which generates the
        ingestion, scoring and LLM-judge programs) and queues creation. Items
        with `choices` make a multiple-choice benchmark; otherwise it is
        question-answer. The benchmark is created unpublished. Requires the
        `benchmark:create` scope. For API tokens and agents, creation is limited
        per user across every creation path (default 10 per rolling 24 hours and
        3 per minute); a dry run does not count, and the web app and platform
        admins are not limited.
      parameters:
        - name: dry_run
          in: query
          required: false
          schema:
            type: boolean
          description: 'Validate only: answers 200 with a summary and creates nothing'
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              required:
                - title
                - items
              properties:
                title:
                  type: string
                  example: Unit Conversion Basics
                description:
                  type: string
                type:
                  type: string
                  enum:
                    - question_answer
                    - multiple_choice
                  description: Optional; inferred from the items when omitted
                items:
                  type: array
                  minItems: 5
                  maxItems: 2000
                  items:
                    type: object
                    required:
                      - question
                      - answer
                    properties:
                      question:
                        type: string
                      answer:
                        type: string
                        description: >-
                          Short text answer, or for choices the letter (A, B,
                          ...) or the exact text of the correct choice
                      choices:
                        type: array
                        items:
                          type: string
                        description: 2-10 options; makes the spec multiple_choice
                      context:
                        type: string
                        description: Optional passage for question_answer items
                      group:
                        type: string
                        description: Optional topic or category used to group results
      responses:
        '200':
          description: 'dry_run: the spec is valid; nothing was created'
        '201':
          description: Creation queued
          content:
            application/json:
              schema:
                type: object
                properties:
                  valid:
                    type: boolean
                  kind:
                    type: string
                  title:
                    type: string
                  status_id:
                    type: integer
                    description: Poll GET /api/competitions/{status_id}/creation_status/
        '400':
          description: Input the caller can fix; `errors` lists every problem
          content:
            application/json:
              schema:
                type: object
                properties:
                  valid:
                    type: boolean
                    example: false
                  errors:
                    type: array
                    items:
                      type: string
                    example:
                      - items[2].answer is empty
        '403':
          description: >-
            Token lacks benchmark:create, or the deployment restricts creation
            to certain roles (`code: role_required`)
        '429':
          description: Creation limit reached; wait for the `Retry-After` header
          headers:
            Retry-After:
              schema:
                type: integer
              description: Seconds until another creation is allowed
          content:
            application/json:
              schema:
                type: object
                properties:
                  code:
                    type: string
                    enum:
                      - quota_exceeded
                      - rate_limited
                  detail:
                    type: string
                  limit:
                    type: integer
                  used:
                    type: integer
                    description: quota_exceeded only
                  resets_at:
                    type: string
                    format: date-time
                    description: quota_exceeded only
                  window_seconds:
                    type: integer
                    description: rate_limited only
      security:
        - platformToken: []
components:
  securitySchemes:
    platformToken:
      type: http
      scheme: bearer
      bearerFormat: bgn_ opaque token
      description: >-
        Platform API token created under Profile Settings > Platform API tokens.
        Scopes are fixed at creation.

````