Fetching the paper…
Reading the bibliography…
Existing benchmarks do not test language agents on their interaction with human users or ability to follow domain-specific rules, both of which are vital for deploying them in real world applications.
Evaluating agenda-based user simulation for reinforcement learning of dialogue management
J. Schatzmann, D. Jurafsky, M. Galley, and D. Trevillian · 2007
Earlier work this paper cites.
Multiwoz–a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
P. Budzianowski, T.-H. Wen, B.-H. Tseng, I. Casanueva, S. Ultes, O. Ramadan, and M. Gašić · 2018
Earlier work this paper cites.
User modeling for task oriented dialogues
I. Gür, D. Hakkani-Tür, G. Tür, and P. Shah · 2018
Earlier work this paper cites.
Decoupling strategy and generation in negotiation dialogues
H. He, D. Chen, A. Balakrishnan, and P. Liang · 2018
Earlier work this paper cites.
Task-oriented dialogue as dataflow synthesis
J. Andreas, J. Bufe, D. Burkett, C. Chen, J. Clausman, J. Crawford, K. Crim, J. DeLoach, L. Dorner, J. Eisner, et al · 2020
Earlier work this paper cites.
Do as I can, not as I say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, et al · 2022
Earlier work this paper cites.
Plm-based world models for text-based games
M. Kim, Y. Jung, D. Lee, and S.-w. Hwang · 2022
Earlier work this paper cites.
ReAct: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2022
Earlier work this paper cites.
Unlocking the potential of user feedback: Leveraging large language model as user simulators to enhance dialogue system
Z. Hu, Y. Feng, A. T. Luu, B. Hooi, and A. Lipani · 2023
Cited alongside, same era.
Metatool benchmark for large language models: Deciding whether to use tools and which to use
Y. Huang, J. Shi, Y. Li, C. Fan, S. Wu, Q. Zhang, Y. Liu, P. Zhou, Y. Wan, N. Z. Gong, et al · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan · 2023
Cited alongside, same era.
Agentbench: Evaluating llms as agents
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, et al · 2023
Cited alongside, same era.
Cognitive architectures for language agents
T. R. Sumers, S. Yao, K. Narasimhan, and T. L. Griffiths · 2023
Later among the works it cites.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Q. Wu, G. Bansal, J. Zhang, Y. Wu, S. Zhang, E. Zhu, B. Li, L. Jiang, X. Zhang, and C. Wang · 2023
Later among the works it cites.
On the tool manipulation capability of open-source large language models, 2023
Q. Xu, F. Hong, B. Li, C. Hu, Z. Chen, and J. Zhang · 2023
Later among the works it cites.
WebArena: A Realistic Web Environment for Building Autonomous Agents
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, Y. Bisk, D. Fried, U. Alon, et al · 2023
Later among the works it cites.
Chatshop: Interactive information seeking with language agents
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein · 2023
Cited alongside, same era.
Identifying the risks of lm agents with an lm-emulated sandbox
Y. Ruan, H. Dong, A. Wang, S. Pitis, Y. Zhou, J. Ba, Y. Dubois, C. J. Maddison, and T. Hashimoto · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Cited alongside, same era.
Reflexion: an autonomous agent with dynamic memory and self-reflection, 2023
N. Shinn, B. Labash, and A. Gopinath · 2023
Cited alongside, same era.
D. Chen, H. Chen, Y. Yang, A. Lin, and Z. Yu
Cited in the paper.
Evaluating large language models trained on code, 2021b
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba
Cited in the paper.
Tree of thoughts: Deliberate problem solving with large language models, 2023a
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan
Cited in the paper.
React: Synergizing reasoning and acting in language models, 2023b
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao
Cited in the paper.
S. Chen, S. Wiseman, and B. Dhingra · 2024
Closest in time.
Evaluating large language models as generative user simulators for conversational recommendation, 2024
S. eun Yoon, Z. He, J. M. Echterhoff, and J. McAuley · 2024
Closest in time.
Berkeley function calling leaderboard
F. Yan, H. Mao, C. C.-J. Ji, T. Zhang, S. G. Patil, I. Stoica, and J. E. Gonzalez · 2024
Closest in time.
Usimagent: Large language models for simulating search users
E. Zhang, X. Wang, P. Gong, Y. Lin, and J. Mao · 2024
Closest in time.