| Rank | Model | Score |
|---|---|---|
| 1 | hy4-preview | 78.6 |
1 phaseActive
Handshake AI Research's benchmark for junior investment-banking agent workflows — data rooms, financial tools, Excel/PowerPoint/Word deliverables. 100 tasks.
Quick answer: BankerToolBench tests AI agents on end-to-end junior investment-banking workflows — navigating data rooms and financial-information tools to produce real deliverables (Excel, PowerPoint, Word, PDF). Built by Handshake AI Research with input from 502 investment bankers, across 100 tasks.
What it tests: Whether an AI agent can complete realistic junior-banker workflows — pulling information from data rooms and financial tools, then producing a polished, correctly-formatted deliverable.
Why it matters: Financial services is one of the highest-value, highest-scrutiny domains for agentic deployment; this measures whether an agent can be trusted with real analyst-level deliverables, not just numeric Q&A.
Known limitations: 100 tasks is a modest sample; grading is rubric-based (weighted criteria) rather than fully automatic, so methodology consistency across evaluators matters.
BankerToolBench was built by Handshake AI Research with direct input from 502 practicing investment bankers, to capture the actual shape of junior-banker work: navigating virtual data rooms, pulling from financial-information tools, and assembling the result into a real office-format deliverable (an Excel model, a PowerPoint deck, a Word memo, or a PDF report). Unlike benchmarks that only check a final numeric answer, BankerToolBench grades the full deliverable against a weighted rubric of expert-defined criteria.
A high score indicates an agent can reliably chain multi-step financial tool-use with correct, professional-grade document production — a capability that's directly relevant to real deployment in financial-services back-office and analyst-support workflows.
| Field | Value |
|---|---|
| Task category | Agent |
| Metric | Weighted rubric score (0–1) |
| Number of tasks | 100 |
| Saturation | Low |
| Created by | Lau, Dücker, Chaudhary et al. (Handshake AI Research) |
| Source paper | Lau et al. 2026 |
| GitHub | Handshake-AI-Research/bankertoolbench |
| Dataset | Hugging Face |
Each task is graded against a set of expert-defined rubric criteria, each binary (met or not met) and weighted 1, 3, 5, or 10 depending on importance. A task's score is the weighted fraction of criteria met, from 0 to 1; the benchmark score is typically reported as this fraction averaged across all tasks (shown here as a percentage).
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Hy4 Preview | 78.6% | Tencent Hunyuan model card | 2026-08 |
Scores sourced from published technical reports and model cards. Results depend on harness, prompt format, and effort settings — see each source for methodology.
No Benchgen results yet — be the first to run BankerToolBench.
| Benchmark | What it tests | Tasks | Saturation |
|---|---|---|---|
| BankerToolBench | Junior investment-banking agent workflows | 100 | Low |
| JobBench | Real job-role task completion, general occupations | — | Low |
| OfficeQA Pro | Office document QA | — | Low |
| Toolathlon-Verified | Multi-tool agentic workflows, general domain | — | Low |
Use BankerToolBench specifically when evaluating an agent for financial-services deployment; use JobBench or Toolathlon-Verified for broader, domain-general tool-use evaluation.
Benchgen lets teams run BankerToolBench against their own model versions, compare results across runs, and catch regressions in financial-workflow reliability — rather than relying on a single vendor-reported number.