Benchgen

BFCL v2 — Results

RankModelScore
1llama-3-3-70b-instruct0.773
2llama-3-1-nemotron-ultra-253b-v10.741
3llama-3-3-nemotron-super-49b-v10.737
4llama-3-2-3b-instruct0.67
5llama-3-1-nemotron-nano-8b-v10.636
B

BFCL v2

1 phaseActive

Extended Berkeley Function Calling Leaderboard with 2,251 enterprise-contributed scenarios, addressing contamination via live user-contributed test cases.

Overview

BFCL v2

Category Metric Tasks Saturation Created

Paper GitHub

Quick answer: BFCL v2 (Patil et al., 2023) extends the Berkeley Function Calling Leaderboard with 2,251 enterprise and OSS-contributed function calling scenarios. It addresses data contamination and bias through live user-contributed test cases across Python, Java, and JavaScript.

What is BFCL v2?

BFCL v2 evaluates AST accuracy, executable accuracy, irrelevance detection, and relevance detection across multiple programming languages with multi-lingual prompts. Enterprise and OSS-contributed functions address real-world function calling scenarios.

Benchmark Details

PropertyValue
Tasks2,251 question-function-answer pairs
MetricOverall accuracy
EvaluationAST + executable accuracy
LanguagesPython, Java, JavaScript

Source: Patil et al. 2023. Last updated 2026-07-24.