Fetching the paper…
Reading the bibliography…
This paper presents a Discriminative Deep Dyna-Q (D3Q) approach to improving the effectiveness and robustness of Deep Dyna-Q (DDQ), a recently proposed framework that extends the Dyna-Q algorithm to integrate planning for task-completion dialogue policy learning.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton. 1990 · 1990
Earlier work this paper cites.
Learning dialogue strategies within the markov decision process framework
Esther Levin, Roberto Pieraccini, and Wieland Eckert. 1997 · 1997
Earlier work this paper cites.
Agenda-based user simulation for bootstrapping a pomdp dialogue system
Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve Young. 2007 · 2007
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky. 2012 · 2012
Earlier work this paper cites.
A survey on metrics for the evaluation of user simulations
Olivier Pietquin and Helen Hastie. 2013 · 2013
Earlier work this paper cites.
Pomdp-based statistical spoken dialog systems: A review
Steve Young, Milica Gašić, Blaise Thomson, and Jason D Williams. 2013 · 2013
Earlier work this paper cites.
Representation learning using multi-task deep neural networks for semantic classification and information retrieval
Xiaodong Liu, Jianfeng Gao, Xiaodong He, Li Deng, Kevin Duh, and Ye-Yi Wang. 2015 · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. 2015 · 2015
Earlier work this paper cites.
Semantically conditioned LSTM-based natural language generation for spoken dialogue systems
Tsung-Hsien Wen, Milica Gasic, Nikola Mrkšić, Pei-Hao Su, David Vandyke, and Steve Young. 2015 · 2015
Earlier work this paper cites.
Continuous deep q-learning with model-based acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine. 2016 · 2016
Earlier work this paper cites.
Multi-domain joint semantic frame parsing using bi-directional rnn-lstm
Dilek Hakkani-Tür, Gokhan Tur, Asli Celikyilmaz, Yun-Nung Chen, Jianfeng Gao, Li Deng, and Ye-Yi Wang. 2016 · 2016
Cited alongside, same era.
A user simulator for task-completion dialogues
Xiujun Li, Zachary C Lipton, Bhuwan Dhingra, Lihong Li, Jianfeng Gao, and Yun-Nung Chen. 2016 · 2016
Cited alongside, same era.
Efficient exploration for dialogue policy learning with bbq networks & replay buffer spiking
Zachary C Lipton, Jianfeng Gao, Lihong Li, Xiujun Li, Faisal Ahmed, and Li Deng. 2016 · 2016
Cited alongside, same era.
The predictron: End-to-end learning and planning
David Silver, Hado van Hasselt, Matteo Hessel, Tom Schaul, Arthur Guez, Tim Harley, Gabriel Dulac-Arnold, David Reichert, Neil Rabinowitz, Andre Barreto, et al. 2016 · 2016
Cited alongside, same era.
End-to-end task-completion neural dialogue systems
Xiujun Li, Yun-Nung Chen, Lihong Li, Jianfeng Gao, and Asli Celikyilmaz. 2017 · 2017
Later among the works it cites.
Iterative policy learning in end-to-end trainable task-oriented neural dialog models
Bing Liu and Ian Lane. 2017 · 2017
Later among the works it cites.
Neural belief tracker: Data-driven dialogue state tracking
Nikola Mrkšić, Diarmuid Ó Séaghdha, Tsung-Hsien Wen, Blaise Thomson, and Steve Young. 2017 · 2017
Later among the works it cites.
Composite task-completion dialogue policy learning via hierarchical deep reinforcement learning
Baolin Peng, Xiujun Li, Lihong Li, Jianfeng Gao, Asli Celikyilmaz, Sungjin Lee, and Kam-Fai Wong. 2017b · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Sébastien Racanière, Théophane Weber, David Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adrià Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pei-Hao Su, Milica Gasic, Nikola Mrksic, Lina Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016 · 2016
Cited alongside, same era.
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel. 2016 · 2016
Cited alongside, same era.
Tiancheng Zhao and Maxine Eskenazi. 2016 · 2016
Cited alongside, same era.
Learning end-to-end goal-oriented dialog
Antoine Bordes, Y-Lan Boureau, and Jason Weston. 2017 · 2017
Cited alongside, same era.
Sub-domain modelling for dialogue management with hierarchical reinforcement learning
Pawel Budzianowski, Stefan Ultes, Pei-Hao Su, Nikola Mrksic, Tsung-Hsien Wen, Inigo Casanueva, Lina Rojas-Barahona, and Milica Gasic. 2017 · 2017
Cited alongside, same era.
Towards end-to-end reinforcement learning of dialogue agents for information access
Bhuwan Dhingra, Lihong Li, Xiujun Li, Jianfeng Gao, Yun-Nung Chen, Faisal Ahmed, and Li Deng. 2017 · 2017
Cited alongside, same era.
Adversarial advantage actor-critic model for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, Yun-Nung Chen, and Kam-Fai Wong. 2017a
Cited in the paper.
A network-based end-to-end trainable task-oriented dialogue system
Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Lina M Rojas-Barahona, Pei-Hao Su, Stefan Ultes, David Vandyke, and Steve Young. 2017 · 2017
Later among the works it cites.
Hybrid code networks: Practical and efficient end-to-end dialog control with supervised and reinforcement learning
Jason D Williams, Kavosh Asadi, and Geoffrey Zweig. 2017 · 2017
Later among the works it cites.
Integrating planning for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, Kam-Fai Wong, and Shang-Yu Su. 2018 · 2018
Closest in time.
Subgoal discovery for hierarchical dialogue policy learning
Da Tang, Xiujun Li, Jianfeng Gao, Chong Wang, Lihong Li, and Tony Jebara. 2018 · 2018
Closest in time.