2023

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Li, Minghao, Zhao, Yingxiu, Yu, Bowen et al.

Understand

Recent research has demonstrated that Large Language Models (LLMs) can enhance their capabilities by utilizing external tools.

  • However, three pivotal questions remain unanswered: (1) How effective are current LLMs in utilizing tools? (2) How can we enhance LLMs' ability to utilize tools? (3) What obstacles need to be overcome to leverage tools? To address these questions, we introduce API-Bank, a groundbreaking benchmark, specifically designed for tool-augmented LLMs.
  • For the first question, we develop a runnable evaluation system consisting of 73 API tools.
  • We annotate 314 tool-use dialogues with 753 API calls to assess the existing LLMs' capabilities in planning, retrieving, and calling APIs.

Reading the bibliography…