Benchgen

FrontierFinance — Results

RankModelScore
1apodex-1-154.3
2apodex-1-1-mini50.2
F

FrontierFinance

1 phaseActive

Samaya Research's agent benchmark for the full investment workflow — 220 expert-rubric-graded queries. Metric: rubric satisfaction rate (%).

Overview

FrontierFinance

Category Metric Tasks Saturation

Paper GitHub Dataset

Quick answer: FrontierFinance is Samaya Research's open benchmark for financial-analyst agent work — 220 expert-crafted queries spanning the full investment workflow (screening, research, financial-model evaluation, event monitoring), each graded against a detailed expert rubric. Apodex 1.1 scores 54.3% as of August 2026, the current best independently-reported model result.

At a Glance

What it tests: An agent's ability to complete real investment-analyst workflows — discovering candidate securities, researching them, extracting and evaluating financial data, and monitoring catalysts/events — not just single-turn financial Q&A.

Why it matters: Most public finance benchmarks (FinanceBench, BigFinanceBench, Vals AI's Finance Agent v2) focus almost entirely on financial data extraction — the easiest sub-task to annotate. FrontierFinance is the first open benchmark to cover the other, harder parts of an analyst's actual job, and its own difficulty analysis shows it has a meaningfully higher mean "hardness" score than any prior public finance benchmark.

Known limitations: Grading relies on an AI-judged expert rubric (not fully automated ground-truth checking), and Samaya Research is both the benchmark's creator and a commercial vendor of financial-research agents — treat vendor-run "system" leaderboard entries (as opposed to raw model scores) with that in mind.

What FrontierFinance Measures

FrontierFinance evaluates AI agents across the full lifecycle of an equity-research workflow, organized into six use cases: Screening & Discovery (finding candidates that meet sector/size/valuation/growth criteria), Company Research, Sector/Industry/Macro analysis, Financial Data Extraction, Coverage & Catalyst tracking, and Earnings & Events monitoring. Each of the 220 queries comes from real analyst work and is paired with an expert-written rubric — an average of 52.5 discrete, independently-checkable rubric items per query (over 11,500 rubric items across the full set), each flagging a specific fact, figure, or claim the answer must correctly surface.

Samaya Research built FrontierFinance specifically to move beyond the "financial data extraction" ceiling that dominates prior finance benchmarks. Its own difficulty analysis pools every example from FinanceBench, BigFinanceBench, and Vals AI's Finance Agent v2 into a single pairwise-compared "hardness" scale (via Bradley-Terry scoring) and finds FrontierFinance has both a higher median and a wider difficulty spread than any of them — a direct consequence of covering harder use cases like coverage/catalyst tracking and screening/discovery rather than pure data lookup.

Benchmark Specifications

FieldValue
Task categoryAgent / finance research
MetricExpert rubric satisfaction rate (%)
Number of queries220
Rubric items~11,500 total (avg. 52.5 per query)
Use cases6 (Screening & Discovery, Company Research, Sector/Industry/Macro, Financial Data Extraction, Coverage & Catalyst, Earnings & Events)
SaturationLow
Created bySamaya Research
PaperarXiv:2608.11683
GitHubsamaya-ai/frontier-finance
Datasetsamaya-ai/FrontierFinance on Hugging Face

How FrontierFinance Is Scored

Each query is graded against its expert-written rubric — a checklist of specific, independently verifiable facts or claims the answer needs to include (some items flagged "must have"). The reported score is the percentage of rubric items satisfied, macro-averaged across all 220 queries. Samaya Research's own published system-level leaderboard (comparing full agent products, not raw models) tops out around 56% as of its latest update — underscoring how much headroom remains versus data-extraction-only benchmarks that are closer to saturated.

State-of-the-Art Results

Scores sourced from published technical reports; Benchgen has not independently re-run these evaluations. Samaya Research's own system-level leaderboard (which pairs models with proprietary agent scaffolds, not raw model results) is not directly comparable to the model-level scores above.

FrontierFinance on Benchgen

No Benchgen results yet — be the first to run FrontierFinance.

FrontierFinance vs Other Benchmarks

BenchmarkWhat it testsTasksSaturation
Finance Agent v2Multi-step financial task automationLow
GDPVal-AA v2Cross-occupation economically valuable workLow
APEX-AgentsProfessional-grade knowledge-work agent tasksLow

FrontierFinance is narrower than GDPVal-AA v2 or APEX-Agents (finance-only), but deeper within its domain — its six-use-case coverage and rubric-based grading go beyond the financial-data-extraction focus of earlier finance-specific benchmarks like Finance Agent v2.

Run FrontierFinance on Your Model

Benchgen lets teams run FrontierFinance against their own model or agent versions, track rubric-level pass rates across releases, and catch regressions in specific investment-workflow use cases — rather than relying on a single vendor-reported headline number.

Frequently Asked Questions

What is FrontierFinance? FrontierFinance is an open benchmark from Samaya Research testing AI agents on the full investment-analyst workflow — screening, research, financial data evaluation, and event monitoring — using 220 queries graded against expert-written rubrics.
What does a good FrontierFinance score look like? As of August 2026, the best independently-reported model score is 54.3% (Apodex 1.1); Samaya Research's own system-level leaderboard tops out around 56%. Scores well below that reflect the benchmark's deliberately broad, difficult use-case coverage.
Who created FrontierFinance? Samaya Research, a financial AI research company, created FrontierFinance and publishes it as an open dataset on Hugging Face alongside a public leaderboard.

Specs from Samaya Research's official benchmark page (research.samaya.ai/benchmarks/frontier-finance) and technical paper (arXiv:2608.11683). Last updated 2026-08-31.