Tool use / function calling
API-Bank
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Heavily superseded#4 of 55 most-superseded · first seen Apr 14, 2023
A heavily-cited baseline — frequently critiqued by newer work
5 papers critique it · 0 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites API-Bank as a baseline.
tested Plan--Retrieve--Call behavior but lacked a unified protocol abstraction
“API-Bank assesses tool-augmented models but limits API candidates to fewer than five per task.”
“Although API-Bank api_bank contains multi-turn interactions, the number of turns in each dialogue is limited (2.84 on average), and the interactions are relatively simplistic.”
“Most public benchmarks still overlook other enterprise-grade challenges, notably distinguishing among near-duplicate tools, proactively eliciting mandatory arguments, and detecting or preventing tool-call hallucinations, shortcomings our framework is expressly designed to remedy.”
“API-Bank~li-etal-2023-api, ToolTalk farn2023tooltalkevaluatingtoolusageconversational and $$-bench yao2024taubenchbenchmarktoolagentuserinteraction does include a set of tools to modify world states, but does not study the impact of state dependencies.”
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.