Fetching the paper…
Reading the bibliography…
Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider.
Evaluating agenda-based user simulation for reinforcement learning of dialogue management
Jost Schatzmann, Daniel Jurafsky, Michael Galley, and David Trevillian · 2007
Earlier work this paper cites.
A survey on metrics for the evaluation of user simulations
Olivier Pietquin and Helen Hastie · 2013
Earlier work this paper cites.
A concise introduction to decentralized POMDPs
Frans A Oliehoek, Christopher Amato, et al · 2016
Earlier work this paper cites.
Multiwoz–a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Inigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić · 2018
Earlier work this paper cites.
User modeling for task oriented dialogues
Izzeddin Gür, Dilek Hakkani-Tür, Gokhan Tür, and Pararth Shah · 2018
Earlier work this paper cites.
Decoupling strategy and generation in negotiation dialogues
He He, Derek Chen, Anusha Balakrishnan, and Percy Liang · 2018
Earlier work this paper cites.
Task-oriented dialogue as dataflow synthesis
Jacob Andreas, John Bufe, David Burkett, Charles Chen, Josh Clausman, Jean Crawford, Kate Crim, Jordan DeLoach, Leah Dorner, Jason Eisner, et al · 2020
Earlier work this paper cites.
Derek Chen, Howard Chen, Yi Yang, Alex Lin, and Zhou Yu · 2021
Earlier work this paper cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan · 2022
Earlier work this paper cites.
Unlocking the potential of user feedback: Leveraging large language model as user simulators to enhance dialogue system
Zhiyuan Hu, Yue Feng, Anh Tuan Luu, Bryan Hooi, and Aldo Lipani · 2023
Earlier work this paper cites.
Metatool benchmark for large language models: Deciding whether to use tools and which to use
Yue Huang, Jiawen Shi, Yuan Li, Chenrui Fan, Siyuan Wu, Qihui Zhang, Yixin Liu, Pan Zhou, Yao Wan, Neil Zhenqiang Gong, et al · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Cited alongside, same era.
Agentbench: Evaluating llms as agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al · 2023
Cited alongside, same era.
Identifying the risks of lm agents with an lm-emulated sandbox
Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J Maddison, and Tatsunori Hashimoto · 2023
Cited alongside, same era.
Berkeley function calling leaderboard
Fanjia Yan, Huanzhi Mao, Charlie Cheng-Jie Ji, Tianjun Zhang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez · 2024
Later among the works it cites.
τ \tau -bench: A benchmark for tool-agent-user interaction in real-world domains
Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan · 2024
Later among the works it cites.
Claude 3.7 Sonnet, 2025
Anthropic · 2025
Closest in time.
litellm, 2025
BerriAI · 2025
Closest in time.
Intellagent: A multi-agent framework for evaluating conversational ai systems
Elad Levi and Ilan Kadar · 2025
Closest in time.
gpt-4.1, 2025
OpenAI · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, et al · 2023
Cited alongside, same era.
Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems, November 2024
Taaha Kazi, Ruiliang Lyu, Sizhe Zhou, Dilek Hakkani-Tur, and Gokhan Tur · 2024
Cited alongside, same era.
Jiarui Lu, Thomas Holleis, Yizhe Zhang, Bernhard Aumayer, Feng Nan, Felix Bai, Shuang Ma, Shen Ma, Mengyu Li, Guoli Yin, et al · 2024
Cited alongside, same era.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al · 2024
Cited alongside, same era.
Flowbench: Revisiting and benchmarking workflow-guided planning for llm-based agents
Ruixuan Xiao, Wentao Ma, Ke Wang, Yuchuan Wu, Junbo Zhao, Haobo Wang, Fei Huang, and Yongbin Li · 2024
Cited alongside, same era.
o4-mini, 2025
OpenAI · 2025
Closest in time.
Apigen-mt: Agentic pipeline for multi-turn data generation via simulated agent-human interplay
Akshara Prabhakar, Zuxin Liu, Weiran Yao, Jianguo Zhang, Ming Zhu, Shiyu Wang, Zhiwei Liu, Tulika Awalgaonkar, Haolin Chen, Thai Hoang, et al · 2025
Closest in time.
Multiagentbench: Evaluating the collaboration and competition of llm agents
Kunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang, Shuyi Guo, Zhe Wang, Zhenhailong Wang, Cheng Qian, Xiangru Tang, Heng Ji, et al · 2025
Closest in time.