Benchgen

Agent Startup Bench — Results

RankModelScore
1seed-2-1-pro0.688
2gpt-5-50.681
3claude-opus-4-70.623
4seed-2-1-turbo0.54
5gemini-3-1-pro0.457
A

Agent Startup Bench

1 phaseActive

ByteDance benchmark measuring AI agents on startup-style tasks requiring autonomous planning and execution. Metric: 0–1 score.

Overview

Agent Startup Bench

Category Metric Saturation Created

Quick answer: Agent Startup Bench is ByteDance's benchmark for evaluating AI agents on high-economic-value, startup-style tasks that demand autonomous multi-step planning, tool use, and execution to deliver practical, verifiable results — simulating the kinds of open-ended workflows that early-stage companies actually run. Seed 2.1 Pro currently leads at 68.8%.

At a Glance

What it tests: An AI agent's ability to autonomously plan and complete realistic startup-style tasks — such as research, analysis, decision-making workflows, and multi-step execution — producing outputs that can be objectively verified.

Why it matters: Unlike benchmarks that test isolated capabilities, Agent Startup Bench evaluates end-to-end agentic behaviour on tasks with real economic stakes, making it directly relevant for teams building or evaluating autonomous agents for business applications.

Known limitations: No public paper, dataset, or evaluation code is available as of mid-2026. Only 2 models are tracked, leaving the score distribution and difficulty calibration poorly characterized.

What Agent Startup Bench Measures

Agent Startup Bench places AI agents in the role of an autonomous contributor at a startup, assigning tasks that require multi-step reasoning, information gathering, planning, and execution. The benchmark emphasises tasks with clear success criteria — outcomes that can be verified objectively — rather than open-ended creative tasks where evaluation is subjective.

The "startup-style" framing reflects tasks that are high-value but under-defined: the agent must decide how to break down a goal, which tools or strategies to use, and when the task is complete. This tests the full autonomy stack rather than isolated capabilities like tool-calling or question answering.

Scores are reported on a 0–1 scale representing the fraction of tasks completed successfully according to the benchmark's verification criteria.

Benchmark Specifications

FieldValue
Task categoryAgent / Autonomous tasks
MetricScore (0–1)
LanguagesEnglish
SaturationLow
Created byByteDance
PaperNot yet public
GitHubNot yet public
DatasetNot yet public

How Agent Startup Bench Is Scored

Each task has a predefined success criterion. The agent is given the task description and optionally a set of tools, and must autonomously produce an output that satisfies the criterion. Scoring is binary per task (success / failure) and the overall score is the fraction of tasks passed, reported on a 0–1 scale.

Agent Startup Bench on Benchgen

No Benchgen results yet — be the first to run Agent Startup Bench.

Agent Startup Bench vs Other Benchmarks

BenchmarkWhat it testsSaturation
Agent Startup BenchStartup-style agentic tasksLow
Tau3 BankingCustomer service agent in bankingLow
SWE-Bench ProSoftware engineering agent tasksLow
MCP AtlasMCP tool-use in agentic workflowsLow

Agent Startup Bench covers a broad, high-value task distribution; use domain-specific agent benchmarks (Tau3, SWE-Bench) when you need targeted coverage of a particular vertical.

Run Agent Startup Bench on Your Model

Benchgen lets you run Agent Startup Bench against your own model versions and agents, compare results across fine-tuning runs, and detect regressions in autonomous task completion before they affect real workflows.

Frequently Asked Questions

What is Agent Startup Bench?

Agent Startup Bench is a ByteDance benchmark that evaluates AI agents on startup-style tasks requiring autonomous planning, tool use, and verifiable execution. It targets high-economic-value, real-world-relevant task scenarios.

What does a good score look like on Agent Startup Bench?

Scores are on a 0–1 scale. As of mid-2026, the top score is 0.688 (Seed 2.1 Pro), with the second model at 0.540 — suggesting the benchmark meaningfully differentiates agent capability levels.

Who created Agent Startup Bench?

Agent Startup Bench was created by ByteDance. No public paper or dataset has been released as of mid-2026.