1 phaseActive
Add adjacent pairs of numbers, then multiply the sums. A verifiable exact-answer benchmark for proving an RLVR/GRPO-trained adapter beats its base model on unseen problems.
Math Pairs: add adjacent pairs of numbers, then multiply the sums.
3, 5, 2, 4 -> 3+5=8, 2+4=6 -> 8 * 6 = 48This is deliberately the same task an RLVR (GRPO) training run optimizes, so the score here is a direct, honest read on whether that training worked.
<answer>...</answer>.<think>...</think>, then only the final number in <answer>...</answer>.The problems are generated from a fixed seed, so every submission is graded on identical problems — that is what makes base-vs-trained a fair comparison.
You do not upload a model.py. Select your model on the platform and it is
called over its OpenAI-compatible /chat/completions endpoint.
On the benchmark's Environment tab, declare three variables with source "From selected model":
| Key | Model field |
|---|---|
RLVR_API_URL | api_url |
RLVR_API_KEY | api_key (secret) |
RLVR_MODEL_NAME | model_name |
Optional tuning variables:
| Key | Meaning | Default |
|---|---|---|
RLVR_TEMPERATURE | keep at 0 (greedy) so scores are reproducible | 0 |
RLVR_MAX_TOKENS | max generated tokens per problem | 256 |
RLVR_WORKERS | parallel requests | 8 |
RLVR_TIMEOUT | per-request timeout (s) | 120 |
If Answered comes back as ~0, these keys are almost certainly not wired.
Run this benchmark twice — once on the RL-trained model and once on its base
model — then compare the Accuracy column on the leaderboard.
Trained > base = the training genuinely improved the model on problems it never saw. That is the one check the in-training reward curve cannot fake.