Fetching the paper…
Reading the bibliography…
Task-oriented dialogue systems (TDSs) are assessed mainly in an offline setting or through human evaluation.
User Modeling for Spoken Dialogue System Evaluation. In IEEE Workshop on ASRU 1997 . 80–87
Wieland Eckert, Esther Levin, and Roberto Pieraccini. 1997 · 1997
Earlier work this paper cites.
PARADISE: A Framework for Evaluating Spoken Dialogue Agents. In Proceedings of the 35th Annual Meeting of the Association for Computational Linguistics and Eighth Conference of the European Chapter of the Association for Computational Linguistics (ACL 1997) . 271–280
Marilyn A. Walker, Diane J. Litman, Candace A. Kamm, and Alicia Abella. 1997 · 1997
Earlier work this paper cites.
The LIMSI ARISE system
Lori Lamel, Sophie Rosset, Jean-Luc Gauvain, Samir Bennacef, Martine Garnier-Rizet, and Bernard Prouts. 2000 · 2000
Earlier work this paper cites.
How to Build User Simulators to Train RL-based Dialog Systems. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP ’19) . 1990–2000
Weiyan Shi, Kun Qian, Xuewei Wang, and Zhou Yu. 2019 · 2000
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL ’02) . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Human-computer Dialogue Simulation using Hidden Markov Models
Heriberto Cuayáhuitl, Steve Renals, Oliver Lemon, and Hiroshi Shimodaira. 2005 · 2005
Earlier work this paper cites.
Learning User Simulations for Information State Update Dialogue Systems. In INTERSPEECH 2005 . 893–896
Kallirroi Georgila, James Henderson, and Oliver Lemon. 2005 · 2005
Earlier work this paper cites.
Dbpedia: A Nucleus for a Web of Open Data
Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. 2007 · 2007
Earlier work this paper cites.
Agenda-Based User Simulation for Bootstrapping a POMDP Dialogue System. In Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; Companion Volume, Short Papers (NAACL ’07) . 149–152
Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve J. Young. 2007 · 2007
Earlier work this paper cites.
Towards Human-like Spoken Dialogue Systems
Jens Edlund, Joakim Gustafson, Mattias Heldner, and Anna Hjalmarsson. 2008 · 2008
Earlier work this paper cites.
Spoken Dialog Challenge 2010: Comparison of Live and Control Test Results. In Proceedings of the SIGDIAL 2011 Conference (SIGDIAL ’11) . 2–7
Alan W. Black, Susanne Burger, Alistair Conkie, H. Hastie, Simon Keizer, Oliver Lemon, Nicolas Merigaud, Gabriel Parent, Gabriel Schubiner, Blaise Thomson, J. Williams, Kai Yu, Steve J. Young, and Maxine Eskénazi. 2011 · 2011
Earlier work this paper cites.
Simulating Simple User Behavior for System Effectiveness Evaluation. In Proceedings of the 20th ACM international conference on Information and knowledge management (CIKM ’11) . 611–620
Ben Carterette, Evangelos Kanoulas, and Emine Yilmaz. 2011 · 2011
Earlier work this paper cites.
Mental Models: an Interdisciplinary Synthesis of Theory and Methods
Natalie A Jones, Helen Ross, Timothy Lynam, Pascal Perez, and Anne Leitch. 2011 · 2011
Earlier work this paper cites.
Real User Evaluation of Spoken Dialogue Systems Using Amazon Mechanical Turk. In INTERSPEECH 2011 . 3061–3064
Filip Jurcícek, Simon Keizer, Milica Gasic, François Mairesse, Blaise Thomson, Kai Yu, and Steve J. Young. 2011 · 2011
Earlier work this paper cites.
Metaphor in Conversation
Anna Kaal. 2012 · 2012
Earlier work this paper cites.
A Survey on Metrics for the Evaluation of User Simulations
Olivier Pietquin and Helen Hastie. 2012 · 2012
Earlier work this paper cites.
POMDP-Based Statistical Spoken Dialog Systems: A Review. In Proceedings of the IEEE , Vol. 101. 1160–1179
Steve Young, Milica Gašić, Blaise Thomson, and Jason D Williams. 2013 · 2013
Earlier work this paper cites.
Multi-domain Dialog State Tracking using Recurrent Neural Networks. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (ACL ’15) . 794–799
Nikola Mrksic, Diarmuid Ó Séaghdha, Blaise Thomson, Milica Gasic, Pei hao Su, David Vandyke, Tsung-Hsien Wen, and Steve J. Young. 2015 · 2015
Earlier work this paper cites.
Oriol Vinyals and Quoc Le. 2015 · 2015
Earlier work this paper cites.
A Diversity-Promoting Objective Function for Neural Conversation Models. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL ’16) . 110–119
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016 · 2016
Earlier work this paper cites.
Agents, Simulated Users and Humans: An Analysis of Performance and Behaviour. In Proceedings of the 25th ACM International on Conference on Information and Knowledge Management (CIKM ’16) . 731––740
David Maxwell and Leif Azzopardi. 2016 · 2016
Earlier work this paper cites.
Neural Belief Tracker: Data-Driven Dialogue State Tracking. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL ’17) . 1777–1788
Nikola Mrksic, Diarmuid Ó Séaghdha, Tsung-Hsien Wen, Blaise Thomson, and Steve J. Young. 2017 · 2017
Earlier work this paper cites.
A Network-based End-to-End Trainable Task-oriented Dialogue System. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers (EACL ’17) . 438–449
Lina Maria Rojas-Barahona, Milica Gasić, Nikola Mrksic, Pei hao Su, Stefan Ultes, Tsung-Hsien Wen, Steve J. Young, and David Vandyke. 2017 · 2017
Earlier work this paper cites.
Information Retrieval Evaluation as Search Simulation: A General Formal Framework for IR Evaluation. In Proceedings of the ACM SIGIR International Conference on Theory of Information Retrieval (ICTIR ’17) . 193–200
Yinan Zhang, Xueqing Liu, and ChengXiang Zhai. 2017 · 2017
Earlier work this paper cites.
Explicit State Tracking with Semi-Supervisionfor Neural Dialogue Generation. In Proceedings of the 27th ACM International Conference on Information and Knowledge Management (CIKM ’18) . 1403–1412
Xisen Jin, Wenqiang Lei, Zhaochun Ren, Hongshen Chen, Shangsong Liang, Yihong Eric Zhao, and Dawei Yin. 2018 · 2018
Cited alongside, same era.
Sequicity: Simplifying Task-oriented Dialogue Systems with Single Sequence-to-Sequence Architectures. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL ’18) . 1437–1447
Wenqiang Lei, Xisen Jin, Min-Yen Kan, Zhaochun Ren, Xiangnan He, and Dawei Yin. 2018 · 2018
Cited alongside, same era.
Towards Deep Conversational Recommendations
Raymond Li, Samira Ebrahimi Kahou, Hannes Schulz, Vincent Michalski, Laurent Charlin, and Christopher Joseph Pal. 2018 · 2018
Cited alongside, same era.
Conversational Recommender System. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval (SIGIR ’18) . 235–244
Yueming Sun and Yi Zhang. 2018 · 2018
Cited alongside, same era.
Conversational AI from an Information Retrieval Perspective: Remaining Challenges and a Case for User Simulation. In Proceedings of the Second International Conference on Design of Experimental Search & Information REtrieval Systems (DESIRES ’21) , Vol. 2950. 80–90
Krisztian Balog. 2021 · 2021
Later among the works it cites.
Sim4IR: The SIGIR 2021 Workshop on Simulation for Information Retrieval Evaluation. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’21) . 2697–2698
Krisztian Balog, David Maxwell, Paul Thomas, and Shuo Zhang. 2021 · 2021
Later among the works it cites.
Advances and Challenges in Conversational Recommender Systems: A Survey
Chongming Gao, Wenqiang Lei, Xiangnan He, Maarten de Rijke, and Tat-Seng Chua. 2021 · 2021
Later among the works it cites.
MultiWOZ 2.3: A Multi-domain Task-Oriented Dialogue Dataset Enhanced with Annotation Corrections and Co-Reference Annotation. In Natural Language Processing and Chinese Computing: 10th CCF International Conference (NLPCC ’21) . 206–218
Ting Han, Ximing Liu, Ryuichi Takanobu, Yixin Lian, Chongxuan Huang, Dazhen Wan, Wei Peng, and Minlie Huang. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards Conversational Search and Recommendation: System Ask, User Respond
Yongfeng Zhang, Xu Chen, Qingyao Ai, Liu Yang, and W. Bruce Croft. 2018 · 2018
Cited alongside, same era.
SIREN: A Simulation Framework for Understanding the Effects of Recommender Systems in Online News Environments. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19) . 150–159
Dimitrios Bountouridis, Jaron Harambam, Mykola Makhortykh, Mónica Marrero, Nava Tintarev, and Claudia Hauff. 2019 · 2019
Cited alongside, same era.
Collaborative Multi-Agent Dialogue Model Training Via Reinforcement Learning. In Proceedings of the 20th Annual SIGdial Meeting on Discourse and Dialogue (SIGdial ’19) . 92–102
Alexandros Papangelis, Yi-Chia Wang, Piero Molino, and Gökhan Tür. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners. In OpenAI blog . 9
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
CoQA: A Conversational Question Answering Challenge
Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
The JDDC Corpus: A Large-Scale Multi-Turn Chinese Dialogue Dataset for E-commerce Customer Service. In Proceedings of the Twelfth Language Resources and Evaluation Conference (LREC ’20) . 459–466
Meng Chen, Ruixue Liu, Lei Shen, Shaozu Yuan, Jingyan Zhou, Youzheng Wu, Xiaodong He, and Bowen Zhou. 2020 · 2020
Cited alongside, same era.
Survey on Evaluation Methods for Dialogue Systems
Jan Deriu, Álvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak. 2020 · 2020
Cited alongside, same era.
MultiWOZ 2.1: A Consolidated Multi-Domain Dialogue Dataset with State Corrections and State Tracking Baselines. In Proceedings of the Twelfth Language Resources and Evaluation Conference (LREC ’20) . 422–428
Mihail Eric, Rahul Goel, Shachi Paul, Adarsh Kumar, Abhishek Sethi, Anuj Kumar Goyal, Peter Ku, Sanchit Agarwal, and Shuyang Gao. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
A Survey on Conversational Recommender Systems
Dietmar Jannach, Ahtsham Manzoor, Wanling Cai, and Li Chen. 2021 · 2021
Later among the works it cites.
An Exploration of Tester-based Evaluation of User Simulators for Comparing Interactive Retrieval Systems. In roceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’21) . 1598–1602
Sahiti Labhishetty and ChengXiang Zhai. 2021 · 2021
Later among the works it cites.
Leveraging Slot Descriptions for Zero-Shot Cross-Domain Dialogue StateTracking. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL ’21) . 5640–5648
Zhaojiang Lin, Bing Liu, Seungwhan Moon, Paul A. Crook, Zhenpeng Zhou, Zhiguang Wang, Zhou Yu, Andrea Madotto, Eunjoon Cho, and Rajen Subba. 2021 · 2021
Later among the works it cites.
CR-Walker: Tree-Structured Graph Reasoning and Dialog Acts for Conversational Recommendation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP ’21) . 1839–1851
Wenchang Ma, Ryuichi Takanobu, and Minlie Huang. 2021 · 2021
Later among the works it cites.
Shades of BLEU, Flavours of Success: The Case of MultiWOZ. In Workshop on GEM 2021 . 34–46
Tomás Nekvinda and Ondrej Dusek. 2021 · 2021
Later among the works it cites.
SOLOIST: Building Task Bots at Scale with Transfer Learning and Machine Teaching
Baolin Peng, Chunyuan Li, Jinchao Li, Shahin Shayandeh, Lars Lidén, and Jianfeng Gao. 2021 · 2021
Later among the works it cites.
KILT: A Benchmark for Knowledge Intensive Language Tasks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL ’21) . 2523–2544
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vassilis Plachouras, Tim Rocktaschel, and Sebastian Riedel. 2021 · 2021
Later among the works it cites.
Wizard of Search Engine: Access to Information Through Conversations with Search Engines. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’21) . 533–543
Pengjie Ren, Zhongkun Liu, Xiaomeng Song, Hongtao Tian, Zhumin Chen, Zhaochun Ren, and Maarten de Rijke. 2021 · 2021
Later among the works it cites.
Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System
Yixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta, Deng Cai, Yi-An Lai, and Yi Zhang. 2021 · 2021
Later among the works it cites.
Simulating User Satisfaction for the Evaluation of Task-oriented Dialogue Systems. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’21) . 2499–2506
Weiwei Sun, Shuo Zhang, Krisztian Balog, Zhaochun Ren, Pengjie Ren, Zhumin Chen, and Maarten de Rijke. 2021 · 2021
Later among the works it cites.
Transferable Dialogue Systems and User Simulators
Bo-Hsiang Tseng, Yinpei Dai, Florian Kreyssig, and Bill Byrne. 2021 · 2021
Later among the works it cites.
UBAR: Towards Fully End-to-End Task-Oriented Dialog Systems with GPT-2. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI ’21) , Vol. 35. 14230–14238
Yunyi Yang, Yunhao Li, and Xiaojun Quan. 2021 · 2021
Later among the works it cites.
Mengzi: Towards Lightweight yet Ingenious Pre-trained Models for Chinese
Zhuosheng Zhang, Hanqing Zhang, Keming Chen, Yuhang Guo, Jingyun Hua, Yulong Wang, and Ming Zhou. 2021 · 2021
Later among the works it cites.
Proceedings of the 1st Workshop on Semiparametric Methods in NLP: Decoupling Logic from Knowledge
Rajarshi Das, Patrick Lewis, Sewon Min, June Thai, and Manzil Zaheer (Eds.). 2022 · 2022
Closest in time.
GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-Supervised Learning and Explicit Policy Injection. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI ’22) , Vol. 36. 10749–10757
Wanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu, Zheng Cao, Dermot Liu, Peng Jiang, Min Yang, Fei Huang, Luo Si, Jian Sun, and Yongbin Li. 2022 · 2022
Closest in time.
A Survey on Security Analysis of Amazon Echo Devices
Surendra Pathak, Sheikh Ariful Islam, Honglu Jiang, Lei Xu, and Emmett Tomai. 2022 · 2022
Closest in time.
Variational Reasoning about User Preferences for Conversational Recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’22) . 165–175
Zhaochun Ren, Zhi Tian, Dongdong Li, Pengjie Ren, Liu Yang, Xin Xin, Huasheng Liang, Maarten de Rijke, and Zhumin Chen. 2022 · 2022
Closest in time.
Eric Michael Smith, Orion Hsu, Rebecca Qian, Stephen Roller, Y-Lan Boureau, and Jason Weston. 2022 · 2022
Closest in time.
Analyzing and Simulating User Utterance Reformulation in Conversational Recommender Systems. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’22) . 133–143
Shuo Zhang, Mu-Chun Wang, and Krisztian Balog. 2022 · 2022
Closest in time.