Benchgen

RLVR Math Pairs (held-out)

1 phaseActive

Add adjacent pairs of numbers, then multiply the sums. A verifiable exact-answer benchmark for proving an RLVR/GRPO-trained adapter beats its base model on unseen problems.

Participate

How to Participate

1. What this benchmark measures

Math Pairs: add adjacent pairs of numbers, then multiply the sums.

3, 5, 2, 4   ->   3+5=8,  2+4=6   ->   8 * 6 = 48

This is deliberately the same task an RLVR (GRPO) training run optimizes, so the score here is a direct, honest read on whether that training worked.

  • Items: a fixed, seeded held-out set (200 problems by default).
  • Metric: exact match on the number inside <answer>...</answer>.
  • Answer format: the model must reply with brief reasoning in <think>...</think>, then only the final number in <answer>...</answer>.

The problems are generated from a fixed seed, so every submission is graded on identical problems — that is what makes base-vs-trained a fair comparison.

2. Your submission is a model endpoint, not code

You do not upload a model.py. Select your model on the platform and it is called over its OpenAI-compatible /chat/completions endpoint.

3. Configure the model connection (required, once)

On the benchmark's Environment tab, declare three variables with source "From selected model":

KeyModel field
RLVR_API_URLapi_url
RLVR_API_KEYapi_key (secret)
RLVR_MODEL_NAMEmodel_name

Optional tuning variables:

KeyMeaningDefault
RLVR_TEMPERATUREkeep at 0 (greedy) so scores are reproducible0
RLVR_MAX_TOKENSmax generated tokens per problem256
RLVR_WORKERSparallel requests8
RLVR_TIMEOUTper-request timeout (s)120

If Answered comes back as ~0, these keys are almost certainly not wired.

4. Prove your training worked

Run this benchmark twice — once on the RL-trained model and once on its base model — then compare the Accuracy column on the leaderboard.

Trained > base = the training genuinely improved the model on problems it never saw. That is the one check the in-training reward curve cannot fake.