| Rank | Model | Score |
|---|---|---|
| 1 | hy4-preview | 39.6 |
1 phaseActive
Compact cross-domain agent-evaluation index built on the Harbor framework, sampling from 6,627 candidate tasks across 54 benchmarks. Metric: pass rate.
Quick answer: Harbor-Index is a compact, cross-domain agent-evaluation index built on the Harbor framework — roughly 80 tasks sampled from a pool of 6,627 candidates across 54 benchmarks spanning software engineering, science, tool use, math, data, and security. Scored as a pass rate.
What it tests: A curated cross-section of agentic capability — a single compact index standing in for a much broader pool of specialist benchmarks.
Why it matters: Running 54 separate benchmarks is expensive; Harbor-Index is designed to give a representative signal of broad agentic capability from a much smaller, official evaluation run.
Known limitations: As a sampled index rather than a from-scratch benchmark, results can shift somewhat across index revisions (task count has varied — 82 tasks at v1.0, 80 in the current Hub revision).
Harbor-Index is built on the Harbor framework — a sandboxed agent-execution environment — and works by selecting a representative subset of tasks from 6,627 candidates spanning 54 existing benchmarks across domains including software engineering, science, tool use, mathematics, data analysis, and security. Rather than introducing entirely new tasks, it functions as a compact "index" whose pass rate is designed to correlate with broad agentic capability across all 54 source benchmarks.
Because it's cross-domain by design, a strong Harbor-Index score suggests general-purpose agentic robustness rather than excellence in any single specialty. The official leaderboard requires running all tasks with at least five trials each, scoring any error as zero — a stricter protocol than a single best-of-N run.
| Field | Value |
|---|---|
| Task category | Agent |
| Metric | Pass rate / accuracy |
| Number of tasks | ~80 (index revision-dependent) |
| Saturation | Low |
| Created by | Harbor-Index team, built on the Harbor framework |
| GitHub | harbor-framework/harbor-index |
| Dataset | Harbor Hub |
Each task returns a binary reward (pass/fail); the index score is the aggregate pass rate across the sampled tasks. The official leaderboard protocol runs all tasks at least five times each and counts any execution error as a failure, reducing variance from lucky single runs.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Hy4 Preview | 39.6% | Tencent Hunyuan model card | 2026-08 |
Scores sourced from published technical reports and model cards. Results depend on harness, prompt format, and effort settings — see each source for methodology.
No Benchgen results yet — be the first to run Harbor-Index.
| Benchmark | What it tests | Tasks | Saturation |
|---|---|---|---|
| Harbor-Index | Cross-domain agentic capability (sampled index) | ~80 | Low |
| Terminal-Bench 2.1 | Long-horizon terminal/shell tasks | — | Low |
| MCP-Atlas | Tool-use via MCP servers | 1000 | Low |
Use Harbor-Index as a single, broad-coverage sanity check on agentic capability before drilling into any one domain-specific benchmark.
Benchgen lets teams run Harbor-Index against their own model versions, compare results across runs, and catch cross-domain regressions — rather than relying on a single vendor-reported number.