Benchgen

τ³ Banking — Results

RankModelScore
1kimi-k333.4
2inkling23.7
3muse-glimmer23.5
4solar-pro-423
5fugu21.7
6fugu-ultra20.6
7nemotron-3-5-lightning-30b-a3b9.5

τ³ Banking

1 phaseActive

Financial domain agent benchmark testing AI on realistic banking workflows — part of the τ³ (tau3) agentic evaluation suite. Metric: % tasks completed.

Overview

τ³ Banking

Category Metric Domain Saturation

Quick answer: τ³ Banking (tau3 Banking) is a financial domain agent benchmark that evaluates AI models on realistic banking and financial service tasks. It is part of the τ³ agentic evaluation suite, which tests models on domain-specific real-world workflows. As of June 2026, Fugu scores 21.7% — matching the top frontier model baselines — while Fugu Ultra scores 20.6%.

At a Glance

What it tests: Domain-specific agentic capability in banking and financial services — completing realistic multi-step tasks such as account management, transaction analysis, and financial queries.

Why it matters: General coding and reasoning benchmarks do not capture how AI performs on specialised industry workflows. τ³ Banking provides a more grounded signal for financial services use cases, where domain knowledge and procedural accuracy both matter.

Known limitations: Financial domain tasks require domain-specific knowledge that may not be well-represented in training data. Scores in the 20–22% range indicate this remains a hard benchmark with substantial room for improvement even at the frontier.

What τ³ Banking Measures

τ³ Banking tests an AI agent's ability to complete structured tasks within banking domain contexts — such as answering customer queries, navigating simulated banking interfaces, and performing financial data analysis. The benchmark simulates realistic back-office and customer-service scenarios, requiring both domain knowledge and reliable instruction-following.

The low absolute scores across all models (20–22% for frontier systems) suggest that financial domain tasks remain genuinely hard, likely due to the combination of domain-specific terminology, procedural requirements, and the need for precise, error-free execution that is characteristic of financial workflows.

Benchmark Specifications

FieldValue
Task categoryAgent / domain-specific
DomainBanking and financial services
Metric% tasks completed
SaturationLow
Suiteτ³ (tau3) agentic evaluation suite

State-of-the-Art Results

RankModelScoreSourceDate
1Fugu21.7%Sakana Fugu technical report2026-06
2Fable 5 / Mythos Preview (max)20.6%Sakana Fugu technical report2026-06
2Fugu Ultra20.6%Sakana Fugu technical report2026-06

Scores sourced from Sakana AI's Fugu technical report, June 2026. The low absolute scores reflect the genuine difficulty of domain-specific financial agentic tasks.