Benchgen

MLS-Bench-Lite — Results

RankModelScore
1kimi-k348.3
2qwen3-8-max41
M

MLS-Bench-Lite

1 phaseActive

Lightweight machine learning systems engineering benchmark testing AI agents on ML infrastructure and pipeline tasks. Metric: % task success.

Overview

MLS-Bench-Lite

Category Metric Saturation Created

Quick answer: MLS-Bench-Lite is a lightweight benchmark testing AI agents on machine learning systems engineering tasks — configuring training pipelines, debugging ML infrastructure, and optimizing model deployment workflows. Kimi K3 scores 48.3% as of July 2026.

At a Glance

What it tests: A model's ability to complete practical ML systems engineering work, such as pipeline configuration and infrastructure debugging, as opposed to model-level ML research tasks.

Why it matters: ML systems engineering (as distinct from ML research) is a large share of real-world ML team workload. MLS-Bench-Lite targets this practical, infrastructure-focused capability.

Known limitations: As a "Lite" variant, task coverage is narrower than a full-scale ML systems benchmark, and independent documentation of methodology is limited.

What MLS-Bench-Lite Measures

MLS-Bench-Lite evaluates AI agents on machine learning systems engineering tasks — the practical infrastructure work surrounding model training and deployment, including pipeline configuration, debugging training failures, and optimizing resource usage. It complements broader coding benchmarks by focusing specifically on ML-adjacent systems work.

Benchmark Specifications

FieldValue
Task categoryCoding / ML systems engineering
Metric% task success
SaturationLow
Created byNot yet independently documented

How MLS-Bench-Lite Is Scored

Agents attempt ML systems engineering tasks, and outcomes are verified against reference configurations or pipeline behavior, producing an aggregate % task success rate.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K348.3%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

MLS-Bench-Lite on Benchgen

No Benchgen results yet — be the first to run MLS-Bench-Lite.

MLS-Bench-Lite vs Other Benchmarks

BenchmarkWhat it testsSaturation
MLS-Bench-LiteML systems engineering tasksLow
PostTrainBenchPost-training/ML research workflowsLow
DeepSWEReal-world software engineering tasksLow
SWE-Bench ProProfessional-grade SWE agent tasksLow

Run MLS-Bench-Lite on Your Model

Benchgen lets you run MLS-Bench-Lite against your own model, tracking ML systems engineering task success as your team evaluates agents for infrastructure and pipeline work.

Frequently Asked Questions

What is MLS-Bench-Lite? MLS-Bench-Lite is a lightweight benchmark testing AI agents on machine learning systems engineering tasks, such as pipeline configuration and infrastructure debugging.
What does a good score look like on MLS-Bench-Lite? Kimi K3 reports 48.3% as of July 2026, reflecting the practical difficulty of real ML systems engineering work for current frontier agents.
Who created MLS-Bench-Lite? MLS-Bench-Lite's originating team is not yet independently documented outside of its citation in Kimi K3's July 2026 technical report.