| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 36.6 |
1 phaseActive
Agentic benchmark testing AI models on post-training and ML research engineering workflows — data curation, fine-tuning, and evaluation pipelines. Metric: % task success.
Quick answer: PostTrainBench is an agentic benchmark testing AI models on post-training and ML research engineering workflows — tasks like curating fine-tuning data, configuring training runs, and building evaluation pipelines. It targets the emerging use case of AI agents assisting with ML research itself. Kimi K3 scores 36.6% as of July 2026.
What it tests: A model's ability to act as an ML research/engineering agent, handling tasks across the post-training lifecycle — data curation, fine-tuning setup, and evaluation.
Why it matters: As labs increasingly use AI agents to accelerate their own research and post-training pipelines, PostTrainBench measures a distinct and practically important capability: can a model help build the next model?
Known limitations: Low scores across the field (Kimi K3's 36.6% is a frontier-level result) indicate this remains an early, unsaturated, and difficult benchmark with limited public documentation of exact task composition.
PostTrainBench evaluates AI agents on realistic ML post-training research and engineering tasks, such as preparing and filtering fine-tuning datasets, configuring RLHF/RLVR training pipelines, and constructing evaluation harnesses. It reflects a meta-capability increasingly relevant to frontier labs: using AI to help accelerate the development of future AI systems.
| Field | Value |
|---|---|
| Task category | Agent / ML research engineering |
| Metric | % task success |
| Saturation | Low |
| Created by | PostTrainBench project |
| Website | posttrainbench.com |
Agents attempt ML research/engineering tasks (e.g., data pipeline construction, training configuration) and are scored on % task success, verified against reference implementations or expected pipeline outputs.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 36.6% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run PostTrainBench.
| Benchmark | What it tests | Saturation |
|---|---|---|
| PostTrainBench | ML post-training research/engineering agent tasks | Low |
| MLS-Bench-Lite | ML systems engineering tasks | Low |
| SWE-Marathon | Long-horizon software engineering | Low |
| DeepSWE | Real-world software engineering tasks | Low |
Benchgen lets you run PostTrainBench against your own model, tracking ML research agent capability over time as your team explores using AI to accelerate model development.