Fetching the paper…
Reading the bibliography…
Deep reinforcement learning is a promising approach to training a dialog manager, but current methods struggle with the large state and action spaces of multi-domain dialog systems.
On the Variance of the Adaptive Learning Rate and Beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han. 2019 · 1908
Earlier work this paper cites.
Guided Dialog Policy Learning: Reward Estimation for Multi-Domain Task-Oriented Dialog
Ryuichi Takanobu, Hanlin Zhu, and Minlie Huang. 2019 · 1908
Earlier work this paper cites.
Towards Scalable Multi-domain Conversational Agents: The Schema-guided Dialogue Dataset
Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. 2019 · 1909
Earlier work this paper cites.
A Plan Recognition Model for Subdialogues in Conversations
Diane J. Litman and James F. Allen. 1987 · 1987
Earlier work this paper cites.
An Application of Reinforcement Learning to Dialogue Strategy Selection in a Spoken Dialogue System for Email
Marilyn A Walker. 2000 · 2000
Earlier work this paper cites.
DIPPER: Description and Formalisation of an Information-State Update Dialogue System Architecture
Johan Bos, Ewan Klein, Oliver Lemon, and Tetsushi Oka. 2003 · 2003
Earlier work this paper cites.
Agenda-Based User Simulation for Bootstrapping a POMDP Dialogue System
Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve Young. 2007 · 2007
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-regret Online Learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell. 2011 · 2011
Earlier work this paper cites.
POMDP-Based Statistical Spoken Dialog Systems: A Review
Steve J. Young, Milica Gasic, Blaise Thomson, and Jason D. Williams. 2013 · 2013
Earlier work this paper cites.
On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
GloVe: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
A Large Annotated Corpus for Learning Natural Language Inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Learning to Search Better Than Your Teacher
Kai-Wei Chang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Hal Daumé III. 2015 · 2015
Earlier work this paper cites.
Human-level Control Through Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. 2015 · 2015
Cited alongside, same era.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. 2015 · 2015
Cited alongside, same era.
Dueling Network Architectures for Deep Reinforcement Learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Van Hasselt, Marc Lanctot, and Nando De Freitas. 2015 · 2015
Cited alongside, same era.
Policy Networks with Two-stage Training for Dialogue Systems
Mehdi Fatemi, Layla El Asri, Hannes Schulz, Jing He, and Kaheer Suleman. 2016 · 2016
Cited alongside, same era.
Playing Hard Exploration Games by Watching YouTube
Yusuf Aytar, Tobias Pfaff, David Budden, Thomas Paine, Ziyu Wang, and Nando de Freitas. 2018 · 2018
Later among the works it cites.
MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Ultes Stefan, Ramadan Osman, and Milica Gašić. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Microsoft Dialogue Challenge: Building End-to-End Task-Completion Dialogue Systems
Xiujun Li, Sarah Panda, Jingjing Liu, and Jianfeng Gao. 2018 · 2018
Later among the works it cites.
BBQ-Networks: Efficient Exploration in Deep Reinforcement Learning for Task-oriented Dialogue Systems
Zachary Lipton, Xiujun Li, Jianfeng Gao, Lihong Li, Faisal Ahmed, and Li Deng. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pei-Hao Su, Milica Gasic, Nikola Mrksic, Lina Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016 · 2016
Cited alongside, same era.
Deep Reinforcement Learning with Double Q-Learning
Hado Van Hasselt, Arthur Guez, and David Silver. 2016 · 2016
Cited alongside, same era.
Frames: A Corpus for Adding Memory to Goal-oriented Dialogue Systems
Layla El Asri, Hannes Schulz, Shikhar Sharma, Jeremie Zumer, Justin Harris, Emery Fine, Rahul Mehrotra, and Kaheer Suleman. 2017 · 2017
Cited alongside, same era.
Learning from Demonstrations for Real World Reinforcement Learning
Todd Hester, Matej Vecerík, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Andrew Sendonaris, Gabriel Dulac-Arnold, Ian Osband, John Agapiou, Joel Z. Leibo, and Audrunas Gruslys. 2017 · 2017
Cited alongside, same era.
Self-Normalizing Neural Networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter. 2017 · 2017
Cited alongside, same era.
End-to-End Task-Completion Neural Dialogue Systems
Xiujun Li, Yun-Nung Chen, Lihong Li, and Jianfeng Gao. 2017 · 2017
Cited alongside, same era.
Iterative Policy Learning in End-to-End Trainable Task-Oriented Neural Dialog Models
Bing Liu and Ian Lane. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Bing Liu, Gokhan Tur, Dilek Hakkani-Tur, Pararth Shah, and Larry Heck. 2018 · 2018
Later among the works it cites.
Learning Montezuma’s Revenge from a Single Demonstration
Tim Salimans and Richard Chen. 2018 · 2018
Later among the works it cites.
Building a Conversational Agent Overnight with Dialogue Self-Play
Pararth Shah, Dilek Hakkani-Tür, Gokhan Tür, Abhinav Rastogi, Ankur Bapna, Neha Nayak, and Larry Heck. 2018 · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Later among the works it cites.
Neural Approaches to Conversational AI
Jianfeng Gao, Michel Galley, Lihong Li, et al. 2019 · 2019
Later among the works it cites.
ConvLab: Multi-Domain End-to-End Dialog System Platform
Sungjin Lee, Qi Zhu, Ryuichi Takanobu, Xiang Li, Yaoqin Zhang, Zheng Zhang, Jinchao Li, Baolin Peng, Xiujun Li, Minlie Huang, and Jianfeng Gao. 2019 · 2019
Later among the works it cites.
Goal-Oriented Dialogue Policy Learning from Failures
Keting Lu, Shiqi Zhang, and Xiaoping Chen. 2019 · 2019
Later among the works it cites.
Show Us the Way: Learning to Manage Dialog from Demonstrations
Gabriel Gordon-Hall, Philip John Gorinski, Gerasimos Lampouras, and Ignacion Iacobacci. 2020 · 2020
Closest in time.