Fetching the paper…
Reading the bibliography…
Traditionally, offline datasets have been used to evaluate task-oriented dialogue (TOD) models.
“Testing the performance of spoken dialogue systems by means of an artificially simulated user,”
Ramón López-Cózar, Zoraida Callejas, and Michael F. McTear, · 2006
Earlier work this paper cites.
“User simulation as testing for spoken dialog systems,”
Hua Ai and Fuliang Weng, · 2008
Earlier work this paper cites.
“A survey on metrics for the evaluation of user simulations,”
Olivier Pietquin and Helen Hastie, · 2013
Earlier work this paper cites.
“A network-based end-to-end trainable task-oriented dialogue system,”
Tsung-Hsien Wen, David Vandyke, Nikola Mrkšić, Milica Gašić, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, and Steve Young, · 2017
Earlier work this paper cites.
“Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems,”
Bing Liu, Gokhan Tür, Dilek Hakkani-Tür, Pararth Shah, and Larry Heck, · 2018
Earlier work this paper cites.
“Improving language understanding by generative pre-training,”
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al., · 2018
Earlier work this paper cites.
“MultiWOZ - a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling,”
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić, · 2018
Earlier work this paper cites.
“Virtualhome: Simulating household activities via programs,”
Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba, · 2018
Earlier work this paper cites.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al., · 2019
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever, · 2019
Earlier work this paper cites.
“Recent advances and challenges in task-oriented dialog systems,”
Zheng Zhang, Ryuichi Takanobu, Qi Zhu, MinLie Huang, and XiaoYan Zhu, · 2020
Earlier work this paper cites.
“Language models are few-shot learners,”
Tom Brown, Benjamin Mann, Nick Ryder, et al., · 2020
Cited alongside, same era.
“Exploring the limits of transfer learning with a unified text-to-text transformer,”
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu, · 2020
Cited alongside, same era.
“BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,”
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer, · 2020
Cited alongside, same era.
“MultiWOZ 2.1: A consolidated multi-domain dialogue dataset with state corrections and state tracking baselines,”
Mihail Eric, Rahul Goel, Shachi Paul, Abhishek Sethi, Sanchit Agarwal, Shuyang Gao, Adarsh Kumar, Anuj Goyal, Peter Ku, and Dilek Hakkani-Tur, · 2020
Cited alongside, same era.
“Teaching models new apis: Domain-agnostic simulators for task oriented dialogue,”
Moya Chen, Paul A. Crook, and Stephen Roller, · 2021
“Llama 2: Open foundation and fine-tuned chat models,” 2023
Hugo Touvron, Louis Martin, Kevin Stone, et al., · 2023
Later among the works it cites.
“User simulation with large language models for evaluating task-oriented dialogue,”
Sam Davidson, Salvatore Romeo, Raphael Shu, James Gung, Arshit Gupta, Saab Mansour, and Yi Zhang, · 2023
Later among the works it cites.
“Metaphorical user simulators for evaluating task-oriented dialogue systems,”
Weiwei Sun, Shuyu Guo, Shuo Zhang, Pengjie Ren, Zhumin Chen, Maarten de Rijke, and Zhaochun Ren, · 2023
Later among the works it cites.
“Holistic evaluation of language models,”
Percy Liang, Rishi Bommasani, and et al., · 2023
Later among the works it cites.
“Judging llm-as-a-judge with mt-bench and chatbot arena,”
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica, · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Measuring massive multitask language understanding,”
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt, · 2021
Cited alongside, same era.
“Is MultiWOZ a solved task? an interactive TOD evaluation framework with user simulator,”
Qinyuan Cheng, Linyang Li, Guofeng Quan, Feng Gao, Xiaofeng Mou, and Xipeng Qiu, · 2022
Cited alongside, same era.
“GenTUS: Simulating user behaviour and language in task-oriented dialogues with generative transformers,”
Hsien-chin Lin, Christian Geishauser, Shutong Feng, Nurul Lubis, Carel van Niekerk, Michael Heck, and Milica Gasic, · 2022
Cited alongside, same era.
“Large language models are zero-shot reasoners,”
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa, · 2022
Cited alongside, same era.
“Do as i can, not as i say: Grounding language in robotic affordances,”
Michael Ahn, Anthony Brohan, Noah Brown, et al., · 2022
Cited alongside, same era.
“Chain-of-thought prompting elicits reasoning in large language models,”
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou, · 2022
Cited alongside, same era.
“Llama: Open and efficient foundation language models,”
Hugo Touvron, Thibaut Lavril, Gautier Izacard, et al., · 2023
Cited alongside, same era.
Later among the works it cites.
“Agentbench: Evaluating llms as agents,”
Xiao Liu, Hao Yu, Hanchen Zhang, et al., · 2023
Later among the works it cites.
“Pptc benchmark: Evaluating large language models for powerpoint task completion,”
Yiduo Guo, Zekai Zhang, Yaobo Liang, Dongyan Zhao, and Duan Nan, · 2023
Later among the works it cites.
“Are llms all you need for task-oriented dialogue?,”
Vojtěch Hudeček and Ondřej Dušek, · 2023
Later among the works it cites.
“Approximating online human evaluation of social chatbots with prompting,”
Ekaterina Svikhnushina and Pearl Pu, · 2023
Later among the works it cites.
“MERCY: Multiple response ranking concurrently in realistic open-domain conversational systems,”
Sarik Ghazarian, Behnam Hedayatnia, Di Jin, Sijia Liu, Nanyun Peng, Yang Liu, and Dilek Hakkani-Tur, · 2023
Later among the works it cites.
“Travelplanner: A benchmark for real-world planning with language agents,”
Jian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu, Renze Lou, Yuandong Tian, Yanghua Xiao, and Yu Su, · 2024
Closest in time.
“Prometheus 2: An open source language model specialized in evaluating other language models,” 2024
Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo, · 2024
Closest in time.