| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 33.4 |
| 2 | inkling | 23.7 |
| 3 | muse-glimmer | 23.5 |
| 4 | solar-pro-4 | 23 |
| 5 | fugu | 21.7 |
| 6 | fugu-ultra | 20.6 |
| 7 | nemotron-3-5-lightning-30b-a3b | 9.5 |
1 phaseActive
Financial domain agent benchmark testing AI on realistic banking workflows — part of the τ³ (tau3) agentic evaluation suite. Metric: % tasks completed.
Quick answer: τ³ Banking (tau3 Banking) is a financial domain agent benchmark that evaluates AI models on realistic banking and financial service tasks. It is part of the τ³ agentic evaluation suite, which tests models on domain-specific real-world workflows. As of June 2026, Fugu scores 21.7% — matching the top frontier model baselines — while Fugu Ultra scores 20.6%.
What it tests: Domain-specific agentic capability in banking and financial services — completing realistic multi-step tasks such as account management, transaction analysis, and financial queries.
Why it matters: General coding and reasoning benchmarks do not capture how AI performs on specialised industry workflows. τ³ Banking provides a more grounded signal for financial services use cases, where domain knowledge and procedural accuracy both matter.
Known limitations: Financial domain tasks require domain-specific knowledge that may not be well-represented in training data. Scores in the 20–22% range indicate this remains a hard benchmark with substantial room for improvement even at the frontier.
τ³ Banking tests an AI agent's ability to complete structured tasks within banking domain contexts — such as answering customer queries, navigating simulated banking interfaces, and performing financial data analysis. The benchmark simulates realistic back-office and customer-service scenarios, requiring both domain knowledge and reliable instruction-following.
The low absolute scores across all models (20–22% for frontier systems) suggest that financial domain tasks remain genuinely hard, likely due to the combination of domain-specific terminology, procedural requirements, and the need for precise, error-free execution that is characteristic of financial workflows.
| Field | Value |
|---|---|
| Task category | Agent / domain-specific |
| Domain | Banking and financial services |
| Metric | % tasks completed |
| Saturation | Low |
| Suite | τ³ (tau3) agentic evaluation suite |
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Fugu | 21.7% | Sakana Fugu technical report | 2026-06 |
| 2 | Fable 5 / Mythos Preview (max) | 20.6% | Sakana Fugu technical report | 2026-06 |
| 2 | Fugu Ultra | 20.6% | Sakana Fugu technical report | 2026-06 |
Scores sourced from Sakana AI's Fugu technical report, June 2026. The low absolute scores reflect the genuine difficulty of domain-specific financial agentic tasks.