Fetching the paper…
Reading the bibliography…
The MultiWOZ dataset (Budzianowski et al.,2018) is frequently used for benchmarking context-to-response abilities of task-oriented dialogue systems.
PARADISE: A Framework for Evaluating Spoken Dialogue Agents
Marilyn A. Walker, Diane J. Litman, Candace A. Kamm, and Alicia Abella. 1997 · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
NLTK: The natural language toolkit
Steven Bird and Edward Loper. 2004 · 2004
Earlier work this paper cites.
A survey of statistical user simulation techniques for reinforcement-learning of dialogue management strategies
Jost Schatzmann, Karl Weilhammer, Matt Stuttle, and Steve Young. 2006 · 2006
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
The Hidden Information State model: A practical framework for POMDP-based spoken dialogue management
Steve Young, Milica Gašić, Simon Keizer, François Mairesse, Jost Schatzmann, Blaise Thomson, and Kai Yu. 2010 · 2010
Earlier work this paper cites.
Learning from real users: rating dialogue success with neural networks for reinforcement learning in spoken dialogue systems
Pei-Hao Su, David Vandyke, Milica Gašić, Dongho Kim, Nikola Mrkšić, Tsung-Hsien Wen, and Steve Young. 2015 · 2011
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
deltaBLEU: A discriminative metric for generation tasks with intrinsically diverse targets
Michel Galley, Chris Brockett, Alessandro Sordoni, Yangfeng Ji, Michael Auli, Chris Quirk, Margaret Mitchell, Jianfeng Gao, and Bill Dolan. 2015 · 2015
Earlier work this paper cites.
Semantically Conditioned LSTM-based Natural Language Generation for Spoken Dialogue Systems
Tsung-Hsien Wen, Milica Gasic, Nikola Mrkšić, Pei-Hao Su, David Vandyke, and Steve Young. 2015 · 2015
Earlier work this paper cites.
Imitation learning for language generation from unaligned data
Gerasimos Lampouras and Andreas Vlachos. 2016 · 2016
Earlier work this paper cites.
A Diversity-Promoting Objective Function for Neural Conversation Models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016 · 2016
Earlier work this paper cites.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
A Copy-Augmented Sequence-to-Sequence Architecture Gives Good Performance on Task-Oriented Dialogue
Mihail Eric and Christopher D. Manning. 2017 · 2017
Earlier work this paper cites.
Why We Need New Evaluation Metrics for NLG
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Earlier work this paper cites.
A network-based end-to-end trainable task-oriented dialogue system
Tsung-Hsien Wen, David Vandyke, Nikola Mrkšić, Milica Gašić, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, and Steve Young. 2017 · 2017
Earlier work this paper cites.
MultiWOZ - a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
Measuring the Diversity of Automatic Image Descriptions
Emiel van Miltenburg, Desmond Elliott, and Piek Vossen. 2018 · 2018
Cited alongside, same era.
Controlling Personality-Based Stylistic Variation with Neural Natural Language Generators
Shereen Oraby, Lena Reed, Shubhangi Tandon, Sharath T.S., Stephanie Lukin, and Marilyn Walker. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Semantically conditioned dialog response generation via hierarchical disentangled self-attention
Wenhu Chen, Jianshu Chen, Pengda Qin, Xifeng Yan, and William Yang Wang. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
MinTL: Minimalist transfer learning for task-oriented dialogue systems
Zhaojiang Lin, Andrea Madotto, Genta Indra Winata, and Pascale Fung. 2020 · 2020
Later among the works it cites.
LAVA: Latent action spaces via variational auto-encoding for dialogue policy optimization
Nurul Lubis, Christian Geishauser, Michael Heck, Hsien-chin Lin, Marco Moresi, Carel van Niekerk, and Milica Gasic. 2020 · 2020
Later among the works it cites.
Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020 · 2020
Later among the works it cites.
USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019 · 2019
Cited alongside, same era.
Evaluating Coherence in Dialogue Systems using Entailment
Nouha Dziri, Ehsan Kamalloo, Kory Mathewson, and Osmar Zaiane. 2019 · 2019
Cited alongside, same era.
Structured fusion networks for dialog
Shikib Mehri, Tejas Srinivasan, and Maxine Eskenazi. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Towards Best Experiment Design for Evaluating Dialogue System Output
Sashank Santhanam and Samira Shaikh. 2019 · 2019
Cited alongside, same era.
Disentangling the Properties of Human Evaluation Methods: A Classification System to Support Comparability, Meta-Evaluation and Reproducibility Testing
Anya Belz, Simon Mille, and David M. Howcroft. 2020 · 2020
Cited alongside, same era.
Evaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge
Ondřej Dušek, Jekaterina Novikova, and Verena Rieser. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Few-shot natural language generation for task-oriented dialog
Baolin Peng, Chenguang Zhu, Chunyuan Li, Xiujun Li, Jinchao Li, Michael Zeng, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Conversation Learner - a machine teaching tool for building dialog managers for task-oriented dialog systems
Swadheen Shukla, Lars Liden, Shahin Shayandeh, Eslam Kamal, Jinchao Li, Matt Mazzola, Thomas Park, Baolin Peng, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
Is Your Goal-Oriented Dialog Model Performing Really Well? Empirical Analysis of System-wise Evaluation
Ryuichi Takanobu, Qi Zhu, Jinchao Li, Baolin Peng, Jianfeng Gao, and Minlie Huang. 2020 · 2020
Later among the works it cites.
Multi-domain dialogue acts and response co-generation
Kai Wang, Junfeng Tian, Rui Wang, Xiaojun Quan, and Jianxing Yu. 2020 · 2020
Later among the works it cites.
MultiWOZ 2.2 : A dialogue dataset with additional annotation corrections and state tracking baselines
Xiaoxue Zang, Abhinav Rastogi, Srinivas Sunkara, Raghav Gupta, Jianguo Zhang, and Jindong Chen. 2020 · 2020
Later among the works it cites.
A probabilistic end-to-end task-oriented dialog model with latent belief states towards semi-supervised learning
Yichi Zhang, Zhijian Ou, Min Hu, and Junlan Feng. 2020a · 2020
Later among the works it cites.
ConvLab-2: An open-source toolkit for building, evaluating, and diagnosing dialogue systems
Qi Zhu, Zheng Zhang, Yan Fang, Xiang Li, Ryuichi Takanobu, Jinchao Li, Baolin Peng, Jianfeng Gao, Xiaoyan Zhu, and Minlie Huang. 2020 · 2020
Later among the works it cites.
Survey on Evaluation Methods for Dialogue Systems
Jan Deriu, Alvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak. 2021 · 2021
Closest in time.
Domain state tracking for a simplified dialogue system
Hyunmin Jeon and Gary Geunbae Lee. 2021 · 2021
Closest in time.
Augpt: Dialogue with pre-trained language models and data augmentation
Jonáš Kulhánek, Vojtěch Hudeček, Tomáš Nekvinda, and Ondřej Dušek. 2021 · 2021
Closest in time.
Modelling hierarchical structure between dialogue policy and natural language generator with option framework for task-oriented dialogue system
Jianhong Wang, Yuan Zhang, Tae-Kyun Kim, and Yunjie Gu. 2021 · 2021
Closest in time.
Alternating recurrent dialog model with large-scale pre-trained language models
Qingyang Wu, Yichi Zhang, Yu Li, and Zhou Yu. 2021 · 2021
Closest in time.