Benchgen

PostTrainBench — Results

RankModelScore
1kimi-k336.6
P

PostTrainBench

1 phaseActive

Agentic benchmark testing AI models on post-training and ML research engineering workflows — data curation, fine-tuning, and evaluation pipelines. Metric: % task success.

Overview

PostTrainBench

Category Metric Saturation Created

Website

Quick answer: PostTrainBench is an agentic benchmark testing AI models on post-training and ML research engineering workflows — tasks like curating fine-tuning data, configuring training runs, and building evaluation pipelines. It targets the emerging use case of AI agents assisting with ML research itself. Kimi K3 scores 36.6% as of July 2026.

At a Glance

What it tests: A model's ability to act as an ML research/engineering agent, handling tasks across the post-training lifecycle — data curation, fine-tuning setup, and evaluation.

Why it matters: As labs increasingly use AI agents to accelerate their own research and post-training pipelines, PostTrainBench measures a distinct and practically important capability: can a model help build the next model?

Known limitations: Low scores across the field (Kimi K3's 36.6% is a frontier-level result) indicate this remains an early, unsaturated, and difficult benchmark with limited public documentation of exact task composition.

What PostTrainBench Measures

PostTrainBench evaluates AI agents on realistic ML post-training research and engineering tasks, such as preparing and filtering fine-tuning datasets, configuring RLHF/RLVR training pipelines, and constructing evaluation harnesses. It reflects a meta-capability increasingly relevant to frontier labs: using AI to help accelerate the development of future AI systems.

Benchmark Specifications

FieldValue
Task categoryAgent / ML research engineering
Metric% task success
SaturationLow
Created byPostTrainBench project
Websiteposttrainbench.com

How PostTrainBench Is Scored

Agents attempt ML research/engineering tasks (e.g., data pipeline construction, training configuration) and are scored on % task success, verified against reference implementations or expected pipeline outputs.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K336.6%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

PostTrainBench on Benchgen

No Benchgen results yet — be the first to run PostTrainBench.

PostTrainBench vs Other Benchmarks

BenchmarkWhat it testsSaturation
PostTrainBenchML post-training research/engineering agent tasksLow
MLS-Bench-LiteML systems engineering tasksLow
SWE-MarathonLong-horizon software engineeringLow
DeepSWEReal-world software engineering tasksLow

Run PostTrainBench on Your Model

Benchgen lets you run PostTrainBench against your own model, tracking ML research agent capability over time as your team explores using AI to accelerate model development.

Frequently Asked Questions

What is PostTrainBench? PostTrainBench is a benchmark testing AI agents on post-training and ML research engineering workflows, including data curation, fine-tuning setup, and evaluation pipeline construction.
What does a good score look like on PostTrainBench? Kimi K3 reports 36.6% as of July 2026. Given the difficulty of genuine ML research engineering tasks, scores in this range represent frontier-level performance.
Who created PostTrainBench? PostTrainBench is maintained as an independent benchmark project (posttrainbench.com) focused on evaluating AI agents on ML post-training research workflows.