Skip to main content
After you train and save a model, you run inference on it by deploying it from its model card. Running a model loads it onto a GPU and exposes an OpenAI-compatible endpoint you can send requests to. This is the same Run Inference flow used for any model on BenchGen, so the full walkthrough lives in Deploy an inference model.

Prerequisites

Steps

1. Open the model and click Run Inference

Open the model’s card and click Run Inference in the top right.
A model card with the Run Inference button outlined in the top right

A model card with the Run Inference button outlined in the top right

2. Configure and start

In the Run Inference dialog, pick a GPU and review the inference settings (max tokens, temperature, top P, context length, precision), then click Run model.
The Run Inference dialog with GPU selection and inference settings, the Run model button outlined

The Run Inference dialog with GPU selection and inference settings, Run model outlined

3. Wait for the endpoint to be ready

The status reads Deploying while the model starts up, then turns to running once it’s live. The model card’s endpoint panel fills in with the model’s identifier, Endpoint URL, and access token.
Logs showing the application is ready and the status set to running

Logs showing the application is ready and the status set to running

For the full deployment walkthrough, including the deployment logs and how to verify the model is serving, see Deploy an inference model.

Send requests

Once the status is running, the endpoint is OpenAI-compatible. Use the Endpoint URL, Token, and model name from the model card:

Next Steps

Monitor Model Usage

Track requests, tokens, latency, and spend for this endpoint.

Evaluate an Inference Model

Benchmark it against an environment and watch the run in real time.

On a Platform Model

Run a formal benchmark to measure improvement.
Last modified on September 3, 2026