Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) methods have emerged as a popular choice for training an efficient and effective dialogue policy.
Convlab: Multi-domain end-to-end dialog system platform
Sungjin Lee, Qi Zhu, Ryuichi Takanobu, Xiang Li, Yaoqin Zhang, Zheng Zhang, Jinchao Li, Baolin Peng, Xiujun Li, Minlie Huang, et al. 2019 · 1904
Earlier work this paper cites.
Statistical theory of extreme values and some practical applications: a series of lectures , volume 33
Emil Julius Gumbel. 1954 · 1954
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Seqgan: Sequence generative adversarial nets with policy gradient
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu. 2017 · 2004
Earlier work this paper cites.
Agenda-based user simulation for bootstrapping a pomdp dialogue system
Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve Young. 2007 · 2007
Earlier work this paper cites.
Partially observable markov decision processes for spoken dialog systems
Jason D Williams and Steve Young. 2007 · 2007
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling. 2013 · 2013
Earlier work this paper cites.
Gaussian processes for pomdp-based dialogue manager optimization
Milica Gašić and Steve Young. 2014 · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015 · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. 2015 · 2015
Cited alongside, same era.
Towards end-to-end reinforcement learning of dialogue agents for information access
Bhuwan Dhingra, Lihong Li, Xiujun Li, Jianfeng Gao, Yun-Nung Chen, Faisal Ahmed, and Li Deng. 2016 · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon. 2016 · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2016 · 2016
Cited alongside, same era.
A user simulator for task-completion dialogues
Xiujun Li, Zachary C Lipton, Bhuwan Dhingra, Lihong Li, Jianfeng Gao, and Yun-Nung Chen. 2016 · 2016
Sample-efficient actor-critic reinforcement learning with supervised data for dialogue management
Pei-Hao Su, Pawel Budzianowski, Stefan Ultes, Milica Gasic, and Steve Young. 2017 · 2017
Later among the works it cites.
Jason D Williams, Kavosh Asadi, and Geoffrey Zweig. 2017 · 2017
Later among the works it cites.
Multiwoz-a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gasic. 2018 · 2018
Later among the works it cites.
Bbq-networks: Efficient exploration in deep reinforcement learning for task-oriented dialogue systems
Zachary Lipton, Xiujun Li, Jianfeng Gao, Lihong Li, Faisal Ahmed, and Li Deng. 2018 · 2018
Later among the works it cites.
Adversarial learning of task-oriented neural dialog models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On-line active reward learning for policy optimisation in spoken dialogue systems
Pei-Hao Su, Milica Gasic, Nikola Mrkšić, Lina M Rojas Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016 · 2016
Cited alongside, same era.
A survey on dialogue systems: Recent advances and new frontiers
Hongshen Chen, Xiaorui Liu, Dawei Yin, and Jiliang Tang. 2017 · 2017
Cited alongside, same era.
End-to-end task-completion neural dialogue systems
Xiujun Li, Yun-Nung Chen, Lihong Li, Jianfeng Gao, and Asli Celikyilmaz. 2017 · 2017
Cited alongside, same era.
Composite task-completion dialogue policy learning via hierarchical deep reinforcement learning
Baolin Peng, Xiujun Li, Lihong Li, Jianfeng Gao, Asli Celikyilmaz, Sungjin Lee, and Kam-Fai Wong. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Cited alongside, same era.
Deep dyna-q: Integrating planning for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Kam-Fai Wong. 2018b
Cited in the paper.
Bing Liu and Ian Lane. 2018 · 2018
Later among the works it cites.
Adversarial advantage actor-critic model for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, Yun-Nung Chen, and Kam-Fai Wong. 2018a · 2018
Later among the works it cites.
Discriminative deep dyna-q: Robust planning for dialogue policy learning
Shang-Yu Su, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Yun-Nung Chen. 2018 · 2018
Later among the works it cites.
Dialogue generation: From imitation learning to inverse reinforcement learning
Ziming Li, Julia Kiseleva, and Maarten de Rijke. 2019 · 2019
Later among the works it cites.
Guided dialog policy learning: Reward estimation for multi-domain task-oriented dialog
Ryuichi Takanobu, Hanlin Zhu, and Minlie Huang. 2019 · 2019
Later among the works it cites.