Fetching the paper…
Reading the bibliography…
This paper presents a new approach that extends Deep Dyna-Q (DDQ) by incorporating a Budget-Conscious Scheduling (BCS) to best utilize a fixed, small amount of user interactions (budget) for learning task-oriented dialogue agents.
Learning dialogue strategies within the markov decision process framework
Esther Levin, Roberto Pieraccini, and Wieland Eckert. 1997 · 1997
Earlier work this paper cites.
Agenda-based user simulation for bootstrapping a pomdp dialogue system
Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve J. Young. 2007 · 2007
Earlier work this paper cites.
The best of both worlds: Unifying conventional dialog systems and pomdps
Jason D Williams. 2008 · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Olivier Chapelle and Lihong Li. 2011 · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton. 2012 · 2012
Earlier work this paper cites.
Pomdp-based statistical spoken dialog systems: A review
Steve J. Young, Milica Gasic, Blaise Thomson, and Jason D. Williams. 2013 · 2013
Earlier work this paper cites.
Thompson sampling with the online bootstrap
Dean Eckles and Maurits Kaptein. 2014 · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. 2015 · 2015
Earlier work this paper cites.
Semantically conditioned lstm-based natural language generation for spoken dialogue systems
Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Pei hao Su, David Vandyke, and Steve J. Young. 2015 · 2015
Earlier work this paper cites.
Policy networks with two-stage training for dialogue systems
Mehdi Fatemi, Layla El Asri, Hannes Schulz, Jing He, and Kaheer Suleman. 2016 · 2016
Earlier work this paper cites.
Multi-domain joint semantic frame parsing using bi-directional rnn-lstm
Dilek Z. Hakkani-Tür, Gökhan Tür, Asli Çelikyilmaz, Yun-Nung Chen, Jianfeng Gao, Li Deng, and Ye-Yi Wang. 2016 · 2016
Cited alongside, same era.
A user simulator for task-completion dialogues
Xiujun Li, Zachary C Lipton, Bhuwan Dhingra, Lihong Li, Jianfeng Gao, and Yun-Nung Chen. 2016 · 2016
Cited alongside, same era.
Continuously learning neural dialogue management
Peihao Su, Milica Gasic, Nikola Mrksic, Lina Maria Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve J. Young. 2016 · 2016
Cited alongside, same era.
Towards end-to-end learning for dialog state tracking and management using deep reinforcement learning
Tiancheng Zhao and Maxine Eskénazi. 2016 · 2016
Cited alongside, same era.
Sub-domain modelling for dialogue management with hierarchical reinforcement learning
Pawel Budzianowski, Stefan Ultes, Pei hao Su, Nikola Mrksic, Tsung-Hsien Wen, Iñigo Casanueva, Lina Maria Rojas-Barahona, and Milica Gasic. 2017 · 2017
Hybrid code networks: practical and efficient end-to-end dialog control with supervised and reinforcement learning
Jason D. Williams, Kavosh Asadi, and Geoffrey Zweig. 2017 · 2017
Later among the works it cites.
Training dialogue systems with human advice
Merwan Barlier, Romain Laroche, and Olivier Pietquin. 2018 · 2018
Later among the works it cites.
Microsoft dialogue challenge: Building end-to-end task-completion dialogue systems
Xiujun Li, Sarah Panda, Jingjing Liu, and Jianfeng Gao. 2018 · 2018
Later among the works it cites.
Efficient exploration for dialogue policy learning with bbq networks & replay buffer spiking
Zachary C Lipton, Jianfeng Gao, Lihong Li, Xiujun Li, Faisal Ahmed, and Li Deng. 2018 · 2018
Later among the works it cites.
Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems
Bing Liu, Gökhan Tür, Dilek Z. Hakkani-Tür, Pararth Shah, and Larry P. Heck. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Affordable on-line dialogue policy learning
Cheng Chang, Runzhe Yang, Lu Chen, Xiang Zhou, and Kai Yu. 2017 · 2017
Cited alongside, same era.
End-to-end reinforcement learning of dialogue agents for information access
Bhuwan Dhingra, Lihong Li, Xiujun Li, Jianfeng Gao, Yun-Nung Chen, Faisal Ahmed, and Li Deng. 2017 · 2017
Cited alongside, same era.
End-to-end task-completion neural dialogue systems
Xiujun Li, Yun-Nung Chen, Lihong Li, and Jianfeng Gao. 2017 · 2017
Cited alongside, same era.
Iterative policy learning in end-to-end trainable task-oriented neural dialog models
Bing Liu and Ian Lane. 2017 · 2017
Cited alongside, same era.
Neural belief tracker: Data-driven dialogue state tracking
Nikola Mrksic, Diarmuid Ó Séaghdha, Tsung-Hsien Wen, Blaise Thomson, and Steve J. Young. 2017 · 2017
Cited alongside, same era.
Integrating planning for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Kam-Fai Wong. 2018 · 2018
Later among the works it cites.
A tutorial on thompson sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, Zheng Wen, et al. 2018 · 2018
Later among the works it cites.
Discriminative deep dyna-q: Robust planning for dialogue policy learning
Shang-Yu Su, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Yun-Nung Chen. 2018 · 2018
Later among the works it cites.
Neural approaches to conversational ai
Jianfeng Gao, Michel Galley, and Lihong Li. 2019 · 2019
Closest in time.
Switch-based active deep dyna-q: Efficient adaptive planning for task-completion dialogue policy learning
Yuexin Wu, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Yiming Yang. 2019 · 2019
Closest in time.