ResearchPod Summary
Function-calling benchmarks are frequently used to rank large language models, yet they are often treated as black boxes that produce a single, reliable scalar score. This study performs a noise floor audit on the Berkeley Function Calling Leaderboard (BFCL) to determine how much of this score is truly stable. By analyzing three endpoints across two providers, the authors distinguish between two types of measurement noise: rerun variability (at a fixed temperature of 0) and prompt-surface sensitivity.
The study finds that reruns at a frozen decoding setting are remarkably stable, with near-deterministic outcomes across the tested endpoints. However, when the researchers introduced semantics-preserving prompt perturbations—such as changing whitespace, adding generic instructions, or wrapping requests in neutral labels—the results shifted significantly. The paired standard deviation for these perturbations was 11 to 58 times larger than the standard deviation observed during simple reruns. This suggests that the "noise floor" of a benchmark is driven more by how a prompt is phrased than by the inherent randomness of the model's output generation.
A critical finding is that marginal accuracy scores can be misleading. The authors developed a failure taxonomy to categorize why tasks fail: structural "malformed-output" errors versus "well-formed but wrong" tool calls. They found that weaker endpoints suffer disproportionately from malformed outputs, while stronger endpoints fail primarily through incorrect arguments or call counts. Because these failure modes are qualitatively different, relying solely on a final success rate hides the structural fragility of the model.
For researchers and developers, this audit highlights that running the same prompt multiple times provides diminishing returns. Instead, evaluation compute is better spent on testing model robustness across a variety of prompt templates and auditing the specific character of failures. The authors propose that leaderboard reporting should include paired perturbation standard deviations and failure-mode breakdowns to provide a more transparent view of model reliability.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.