| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 48.3 |
| 2 | qwen3-8-max | 41 |
1 phaseActive
Lightweight machine learning systems engineering benchmark testing AI agents on ML infrastructure and pipeline tasks. Metric: % task success.
Quick answer: MLS-Bench-Lite is a lightweight benchmark testing AI agents on machine learning systems engineering tasks — configuring training pipelines, debugging ML infrastructure, and optimizing model deployment workflows. Kimi K3 scores 48.3% as of July 2026.
What it tests: A model's ability to complete practical ML systems engineering work, such as pipeline configuration and infrastructure debugging, as opposed to model-level ML research tasks.
Why it matters: ML systems engineering (as distinct from ML research) is a large share of real-world ML team workload. MLS-Bench-Lite targets this practical, infrastructure-focused capability.
Known limitations: As a "Lite" variant, task coverage is narrower than a full-scale ML systems benchmark, and independent documentation of methodology is limited.
MLS-Bench-Lite evaluates AI agents on machine learning systems engineering tasks — the practical infrastructure work surrounding model training and deployment, including pipeline configuration, debugging training failures, and optimizing resource usage. It complements broader coding benchmarks by focusing specifically on ML-adjacent systems work.
| Field | Value |
|---|---|
| Task category | Coding / ML systems engineering |
| Metric | % task success |
| Saturation | Low |
| Created by | Not yet independently documented |
Agents attempt ML systems engineering tasks, and outcomes are verified against reference configurations or pipeline behavior, producing an aggregate % task success rate.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 48.3% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run MLS-Bench-Lite.
| Benchmark | What it tests | Saturation |
|---|---|---|
| MLS-Bench-Lite | ML systems engineering tasks | Low |
| PostTrainBench | Post-training/ML research workflows | Low |
| DeepSWE | Real-world software engineering tasks | Low |
| SWE-Bench Pro | Professional-grade SWE agent tasks | Low |
Benchgen lets you run MLS-Bench-Lite against your own model, tracking ML systems engineering task success as your team evaluates agents for infrastructure and pipeline work.