Fetching the paper…
Reading the bibliography…
Training a task-completion dialogue agent via reinforcement learning (RL) is costly because it requires many interactions with real users.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton. 1990 · 1990
Earlier work this paper cites.
Reinforcement learning with a hierarchy of abstract models
Satinder P Singh. 1992 · 1992
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Andrew W Moore and Christopher G Atkeson. 1993 · 1993
Earlier work this paper cites.
Efficient learning and planning within the dyna framework
Jing Peng and Ronald J Williams. 1993 · 1993
Earlier work this paper cites.
Model-based reinforcement learning with an approximate, learned model
Leonid Kuvayev and Richard S Sutton. 1996 · 1996
Earlier work this paper cites.
Learning dialogue strategies within the markov decision process framework
Esther Levin, Roberto Pieraccini, and Wieland Eckert. 1997 · 1997
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton and Andrew G Barto. 1998 · 1998
Earlier work this paper cites.
Optimizing dialogue management with reinforcement learning: Experiments with the njfun system
Satinder Singh, Diane Litman, Michael Kearns, and Marilyn Walker. 2002 · 2002
Earlier work this paper cites.
Agenda-based user simulation for bootstrapping a pomdp dialogue system
Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve Young. 2007 · 2007
Earlier work this paper cites.
Gaussian processes for fast policy optimisation of pomdp-based dialogue managers
Milica Gašić, Filip Jurčíček, Simon Keizer, François Mairesse, Blaise Thomson, Kai Yu, and Steve Young. 2010 · 2010
Earlier work this paper cites.
On-line policy optimisation of spoken dialogue systems via live interaction with human subjects
Milica Gašić, Filip Jurčíček, Blaise Thomson, Kai Yu, and Steve Young. 2011 · 2011
Earlier work this paper cites.
Sample efficient on-line learning of optimal dialogue policies with kalman temporal differences
Olivier Pietquin, Matthieu Geist, Senthilkumar Chandramohan, et al. 2011 · 2011
Cited alongside, same era.
A comparative study of reinforcement learning techniques on dialogue management
Alexandros Papangelis. 2012 · 2012
Cited alongside, same era.
Dyna-style planning with linear function approximation and prioritized sweeping
Richard S Sutton, Csaba Szepesvári, Alborz Geramifard, and Michael P Bowling. 2012 · 2012
Cited alongside, same era.
A survey on metrics for the evaluation of user simulations
Olivier Pietquin and Helen Hastie. 2013 · 2013
Cited alongside, same era.
Pomdp-based statistical spoken dialog systems: A review
Steve Young, Milica Gašić, Blaise Thomson, and Jason D Williams. 2013 · 2013
Cited alongside, same era.
Neural belief tracker: Data-driven dialogue state tracking
Nikola Mrkšić, Diarmuid O Séaghdha, Tsung-Hsien Wen, Blaise Thomson, and Steve Young. 2016 · 2016
Later among the works it cites.
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel. 2016 · 2016
Later among the works it cites.
Tiancheng Zhao and Maxine Eskenazi. 2016 · 2016
Later among the works it cites.
Sub-domain modelling for dialogue management with hierarchical reinforcement learning
Pawel Budzianowski, Stefan Ultes, Pei-Hao Su, Nikola Mrksic, Tsung-Hsien Wen, Inigo Casanueva, Lina Rojas-Barahona, and Milica Gasic. 2017 · 2017
Later among the works it cites.
Towards end-to-end reinforcement learning of dialogue agents for information access
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Representation learning using multi-task deep neural networks for semantic classification and information retrieval
Xiaodong Liu, Jianfeng Gao, Xiaodong He, Li Deng, Kevin Duh, and Ye-Yi Wang. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. 2015 · 2015
Cited alongside, same era.
Semantically conditioned lstm-based natural language generation for spoken dialogue systems
Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Pei-Hao Su, David Vandyke, and Steve Young. 2015 · 2015
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine. 2016 · 2016
Cited alongside, same era.
Multi-domain joint semantic frame parsing using bi-directional RNN-LSTM
Dilek Hakkani-Tür, Gokhan Tur, Asli Celikyilmaz, Yun-Nung Chen, Jianfeng Gao, Li Deng, and Ye-Yi Wang. 2016 · 2016
Cited alongside, same era.
Efficient exploration for dialogue policy learning with bbq networks & replay buffer spiking
Zachary C Lipton, Jianfeng Gao, Lihong Li, Xiujun Li, Faisal Ahmed, and Li Deng. 2016 · 2016
Cited alongside, same era.
Dialogue learning with human-in-the-loop
Jiwei Li, Alexander H Miller, Sumit Chopra, Marc’Aurelio Ranzato, and Jason Weston. 2016a
Cited in the paper.
Bhuwan Dhingra, Lihong Li, Xiujun Li, Jianfeng Gao, Yun-Nung Chen, Faisal Ahmed, and Li Deng. 2017 · 2017
Later among the works it cites.
End-to-end task-completion neural dialogue systems
Xuijun Li, Yun-Nung Chen, Lihong Li, Jianfeng Gao, and Asli Celikyilmaz. 2017 · 2017
Later among the works it cites.
Iterative policy learning in end-to-end trainable task-oriented neural dialog models
Bing Liu and Ian Lane. 2017 · 2017
Later among the works it cites.
Composite task-completion dialogue policy learning via hierarchical deep reinforcement learning
Baolin Peng, Xiujun Li, Lihong Li, Jianfeng Gao, Asli Celikyilmaz, Sungjin Lee, and Kam-Fai Wong. 2017b · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Sébastien Racanière, Théophane Weber, David Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adrià Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al. 2017 · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al. 2017 · 2017
Later among the works it cites.
Hybrid code networks: Practical and efficient end-to-end dialog control with supervised and reinforcement learning
Jason D Williams, Kavosh Asadi, and Geoffrey Zweig. 2017 · 2017
Later among the works it cites.