Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have exhibited the ability to effectively utilize external tools to address user queries.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Earlier work this paper cites.
Learning by distilling context
C. Snell, D. Klein, and R. Zhong · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Gorilla: Large language model connected with massive apis
S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao · 2023
Earlier work this paper cites.
Granite-function calling model: Introducing function calling abilities via multi-task learning of granular tasks
I. Abdelaziz, K. Basu, M. Agarwal, S. Kumaravel, M. Stallone, R. Panda, Y. Rizk, G. P. S. Bhargav, M. Crouse, C. Gunasekara, S. Ikbal, S. Joshi, H. Karanam, V. Kumar, A. Munawar, S. Neelam, D. Raghu, U. Sharma, A. M. Soria, D. Sreedhar, P. Venkateswaran, M. Unuvar, D. D. Cox, S. Roukos, L. A. Lastras, and P. Kapanipathi · 2024
Earlier work this paper cites.
Nestful: A benchmark for evaluating llms on nested sequences of api calls
K. Basu, I. Abdelaziz, K. Bradford, M. Crouse, K. Kate, S. Kumaravel, S. Goyal, A. Munawar, Y. Rizk, X. Wang, et al · 2024
Earlier work this paper cites.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Cited alongside, same era.
Stabletoolbench: Towards stable large-scale benchmarking on tool learning of large language models
Z. Guo, S. Cheng, H. Wang, S. Liang, Y. Qin, P. Li, Z. Liu, M. Sun, and Y. Liu · 2024
Cited alongside, same era.
Qwen2. 5-coder technical report
B. Hui, J. Yang, Z. Cui, J. Yang, D. Liu, L. Zhang, T. Liu, J. Zhang, B. Yu, K. Lu, et al · 2024
Cited alongside, same era.
Self-alignment with instruction backtranslation
X. Li, P. Yu, C. Zhou, T. Schick, O. Levy, L. Zettlemoyer, J. E. Weston, and M. Lewis · 2024
Cited alongside, same era.
Hammer: Robust function-calling for on-device language models via function masking
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
G. Team, P. Georgiev, V. I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang, et al · 2024
Later among the works it cites.
Building math agents with multi-turn iterative preference learning
W. Xiong, C. Shi, J. Shen, A. Rosenberg, Z. Qin, D. Calandriello, M. Khalman, R. Joshi, B. Piot, M. Saleh, et al · 2024
Later among the works it cites.
Berkeley function calling leaderboard
F. Yan, H. Mao, C. C.-J. Ji, T. Zhang, S. G. Patil, I. Stoica, and J. E. Gonzalez · 2024
Later among the works it cites.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al · 2024
Later among the works it cites.
tau-bench: A benchmark for tool-agent-user interaction in real-world domains
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Q. Lin, M. Wen, Q. Peng, G. Nie, J. Liao, J. Wang, X. Mo, J. Zhou, C. Cheng, Y. Zhao, et al · 2024
Cited alongside, same era.
J. Lu, T. Holleis, Y. Zhang, B. Aumayer, F. Nan, F. Bai, S. Ma, S. Ma, M. Li, G. Yin, et al · 2024
Cited alongside, same era.
Agentboard: An analytical evaluation board of multi-turn llm agents
C. Ma, J. Zhang, Z. Zhu, C. Yang, Y. Yang, Y. Jin, Z. Lan, L. Kong, and J. He · 2024
Cited alongside, same era.
Better alignment with instruction back-and-forth translation
T. Nguyen, J. Li, S. Oh, L. Schmidt, J. E. Weston, L. Zettlemoyer, and X. Li · 2024
Cited alongside, same era.
ToolLLM: Facilitating large language models to master 16000+ real-world APIs
Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, S. Zhao, L. Hong, R. Tian, R. Xie, J. Zhou, M. Gerstein, dahai li, Z. Liu, and M. Sun · 2024
Cited alongside, same era.
Facilitating multi-turn function calling for llms via compositional instruction tuning
M. Chen, H. Sun, T. Li, F. Yang, H. Liang, K. Lu, B. Cui, W. Zhang, Z. Zhou, and W. Chen
Cited in the paper.
Re-invoke: Tool invocation rewriting for zero-shot tool retrieval
Y. Chen, J. Yoon, D. S. Sachan, Q. Wang, V. Cohen-Addad, M. Bateni, C.-Y. Lee, and T. Pfister
Cited in the paper.
Toolace: Winning the points of llm function calling
W. Liu, X. Huang, X. Zeng, X. Hao, S. Yu, D. Li, S. Wang, W. Gan, Z. Liu, Y. Yu, et al
Cited in the paper.
S. Yao, N. Shinn, P. Razavi, and K. Narasimhan · 2024
Later among the works it cites.
Agent lumos: Unified and modular training for open-source language agents
D. Yin, F. Brahman, A. Ravichander, K. Chandu, K.-W. Chang, Y. Choi, and B. Y. Lin · 2024
Later among the works it cites.
xlam: A family of large action models to empower ai agent systems
J. Zhang, T. Lan, M. Zhu, Z. Liu, T. Hoang, S. Kokane, W. Yao, J. Tan, A. Prabhakar, H. Chen, et al · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.