Benchgen

Harbor-Index — Results

RankModelScore
1hy4-preview39.6
H

Harbor-Index

1 phaseActive

Compact cross-domain agent-evaluation index built on the Harbor framework, sampling from 6,627 candidate tasks across 54 benchmarks. Metric: pass rate.

Overview

Harbor-Index

Category Metric Tasks Saturation Created

GitHub

Quick answer: Harbor-Index is a compact, cross-domain agent-evaluation index built on the Harbor framework — roughly 80 tasks sampled from a pool of 6,627 candidates across 54 benchmarks spanning software engineering, science, tool use, math, data, and security. Scored as a pass rate.

At a Glance

What it tests: A curated cross-section of agentic capability — a single compact index standing in for a much broader pool of specialist benchmarks.

Why it matters: Running 54 separate benchmarks is expensive; Harbor-Index is designed to give a representative signal of broad agentic capability from a much smaller, official evaluation run.

Known limitations: As a sampled index rather than a from-scratch benchmark, results can shift somewhat across index revisions (task count has varied — 82 tasks at v1.0, 80 in the current Hub revision).

What Harbor-Index Measures

Harbor-Index is built on the Harbor framework — a sandboxed agent-execution environment — and works by selecting a representative subset of tasks from 6,627 candidates spanning 54 existing benchmarks across domains including software engineering, science, tool use, mathematics, data analysis, and security. Rather than introducing entirely new tasks, it functions as a compact "index" whose pass rate is designed to correlate with broad agentic capability across all 54 source benchmarks.

Because it's cross-domain by design, a strong Harbor-Index score suggests general-purpose agentic robustness rather than excellence in any single specialty. The official leaderboard requires running all tasks with at least five trials each, scoring any error as zero — a stricter protocol than a single best-of-N run.

Benchmark Specifications

FieldValue
Task categoryAgent
MetricPass rate / accuracy
Number of tasks~80 (index revision-dependent)
SaturationLow
Created byHarbor-Index team, built on the Harbor framework
GitHubharbor-framework/harbor-index
DatasetHarbor Hub

How Harbor-Index Is Scored

Each task returns a binary reward (pass/fail); the index score is the aggregate pass rate across the sampled tasks. The official leaderboard protocol runs all tasks at least five times each and counts any execution error as a failure, reducing variance from lucky single runs.

State-of-the-Art Results

RankModelScoreSourceDate
1Hy4 Preview39.6%Tencent Hunyuan model card2026-08

Scores sourced from published technical reports and model cards. Results depend on harness, prompt format, and effort settings — see each source for methodology.

Harbor-Index on Benchgen

No Benchgen results yet — be the first to run Harbor-Index.

Harbor-Index vs Other Benchmarks

BenchmarkWhat it testsTasksSaturation
Harbor-IndexCross-domain agentic capability (sampled index)~80Low
Terminal-Bench 2.1Long-horizon terminal/shell tasksLow
MCP-AtlasTool-use via MCP servers1000Low

Use Harbor-Index as a single, broad-coverage sanity check on agentic capability before drilling into any one domain-specific benchmark.

Run Harbor-Index on Your Model

Benchgen lets teams run Harbor-Index against their own model versions, compare results across runs, and catch cross-domain regressions — rather than relying on a single vendor-reported number.

Frequently Asked Questions

What is Harbor-Index? Harbor-Index is a compact cross-domain agent-evaluation index built on the Harbor framework, sampling roughly 80 tasks from 6,627 candidates across 54 benchmarks.
What does a good Harbor-Index score look like? As of August 2026, frontier models score in the 35–47% range; scores above 40% represent strong broad-domain agentic capability.
Who created Harbor-Index? Harbor-Index was created by the Harbor-Index team (Lin Shi, Haowei Lin, Zixuan Zhu, Xiaoyue Zhou, Xiang Li), built on the Harbor framework.