| Rank | Model | Score |
|---|---|---|
| 1 | seed-2-1-pro | 0.688 |
| 2 | gpt-5-5 | 0.681 |
| 3 | claude-opus-4-7 | 0.623 |
| 4 | seed-2-1-turbo | 0.54 |
| 5 | gemini-3-1-pro | 0.457 |
1 phaseActive
ByteDance benchmark measuring AI agents on startup-style tasks requiring autonomous planning and execution. Metric: 0–1 score.
Quick answer: Agent Startup Bench is ByteDance's benchmark for evaluating AI agents on high-economic-value, startup-style tasks that demand autonomous multi-step planning, tool use, and execution to deliver practical, verifiable results — simulating the kinds of open-ended workflows that early-stage companies actually run. Seed 2.1 Pro currently leads at 68.8%.
What it tests: An AI agent's ability to autonomously plan and complete realistic startup-style tasks — such as research, analysis, decision-making workflows, and multi-step execution — producing outputs that can be objectively verified.
Why it matters: Unlike benchmarks that test isolated capabilities, Agent Startup Bench evaluates end-to-end agentic behaviour on tasks with real economic stakes, making it directly relevant for teams building or evaluating autonomous agents for business applications.
Known limitations: No public paper, dataset, or evaluation code is available as of mid-2026. Only 2 models are tracked, leaving the score distribution and difficulty calibration poorly characterized.
Agent Startup Bench places AI agents in the role of an autonomous contributor at a startup, assigning tasks that require multi-step reasoning, information gathering, planning, and execution. The benchmark emphasises tasks with clear success criteria — outcomes that can be verified objectively — rather than open-ended creative tasks where evaluation is subjective.
The "startup-style" framing reflects tasks that are high-value but under-defined: the agent must decide how to break down a goal, which tools or strategies to use, and when the task is complete. This tests the full autonomy stack rather than isolated capabilities like tool-calling or question answering.
Scores are reported on a 0–1 scale representing the fraction of tasks completed successfully according to the benchmark's verification criteria.
| Field | Value |
|---|---|
| Task category | Agent / Autonomous tasks |
| Metric | Score (0–1) |
| Languages | English |
| Saturation | Low |
| Created by | ByteDance |
| Paper | Not yet public |
| GitHub | Not yet public |
| Dataset | Not yet public |
Each task has a predefined success criterion. The agent is given the task description and optionally a set of tools, and must autonomously produce an output that satisfies the criterion. Scoring is binary per task (success / failure) and the overall score is the fraction of tasks passed, reported on a 0–1 scale.
No Benchgen results yet — be the first to run Agent Startup Bench.
| Benchmark | What it tests | Saturation |
|---|---|---|
| Agent Startup Bench | Startup-style agentic tasks | Low |
| Tau3 Banking | Customer service agent in banking | Low |
| SWE-Bench Pro | Software engineering agent tasks | Low |
| MCP Atlas | MCP tool-use in agentic workflows | Low |
Agent Startup Bench covers a broad, high-value task distribution; use domain-specific agent benchmarks (Tau3, SWE-Bench) when you need targeted coverage of a particular vertical.
Benchgen lets you run Agent Startup Bench against your own model versions and agents, compare results across fine-tuning runs, and detect regressions in autonomous task completion before they affect real workflows.
Agent Startup Bench is a ByteDance benchmark that evaluates AI agents on startup-style tasks requiring autonomous planning, tool use, and verifiable execution. It targets high-economic-value, real-world-relevant task scenarios.
Scores are on a 0–1 scale. As of mid-2026, the top score is 0.688 (Seed 2.1 Pro), with the second model at 0.540 — suggesting the benchmark meaningfully differentiates agent capability levels.
Agent Startup Bench was created by ByteDance. No public paper or dataset has been released as of mid-2026.