Fetching the paper…
Reading the bibliography…
Human conversation is inherently complex, often spanning many different topics/domains.
Learning and executing generalized robot plans
Richard E Fikes, Peter E Hart, and Nils J Nilsson. 1972 · 1972
Earlier work this paper cites.
Chunking in soar: The anatomy of a general learning mechanism
John E Laird, Paul S Rosenbloom, and Allen Newell. 1986 · 1986
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Ronald Parr and Stuart J Russell. 1998 · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto. 1999 · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich. 2000 · 2000
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Andrew G Barto and Sridhar Mahadevan. 2003 · 2003
Earlier work this paper cites.
A survey of statistical user simulation techniques for reinforcement-learning of dialogue management strategies
Jost Schatzmann, Karl Weilhammer, Matt Stuttle, and Steve Young. 2006 · 2006
Earlier work this paper cites.
Partially observable Markov decision processes for spoken dialog systems
Jason D. Williams and Steve Young. 2007 · 2007
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
Hierarchical reinforcement learning for spoken dialogue systems
Heriberto Cuayáhuitl. 2009 · 2009
Earlier work this paper cites.
Evaluation of a hierarchical reinforcement learning spoken dialogue system
Heriberto Cuayáhuitl, Steve Renals, Oliver Lemon, and Hiroshi Shimodaira. 2010 · 2010
Cited alongside, same era.
On-line policy optimisation of spoken dialogue systems via live interaction with human subjects
Milica Gašić, Filip Jurcicek, Blaise. Thomson, Kai Yu, and Steve Young. 2011 · 2011
Cited alongside, same era.
Multi-policy dialogue management
Pierre Lison. 2011 · 2011
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013 · 2013
Cited alongside, same era.
Pomdp-based statistical spoken dialog systems: A review
Steve Young, Milica Gašić, Blaise Thomson, and Jason D Williams. 2013 · 2013
Cited alongside, same era.
Deep reinforcement learning for multi-domain dialogue systems
Heriberto Cuayáhuitl, Seunghak Yu, Ashley Williamson, and Jacob Carse. 2016 · 2016
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel. 2016 · 2016
Later among the works it cites.
Policy networks with two-stage training for dialogue systems
Mehdi Fatemi, Layla El Asri, Hannes Schulz, Jing He, and Kaheer Suleman. 2016 · 2016
Later among the works it cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum. 2016 · 2016
Later among the works it cites.
Continuously learning neural dialogue management
Pei-Hao Su, Milica Gasic, Nikola Mrksic, Lina Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016 · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Incremental on-line adaptation of pomdp-based dialogue managers to extended domains
Milica Gašić, Dongho Kim, Pirros Tsiakoulis, Catherine Breslin, Matthew Henderson, Martin Szummer, Blaise Thomson, and Steve Young. 2014 · 2014
Cited alongside, same era.
Gaussian processes for pomdp-based dialogue manager optimization
Milica Gašić and Steve Young. 2014 · 2014
Cited alongside, same era.
Policy learning for domain selection in an extensible multi-domain spoken dialogue system
Zhuoran Wang, Hongliang Chen, Guanchun Wang, Hao Tian, Hua Wu, and Haifeng Wang. 2014 · 2014
Cited alongside, same era.
Policy committee for adaptation in multi-domain spoken dialogue systems
Milica Gašić, Nikola Mrkšić, Pei-hao Su, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2015 · 2015
Cited alongside, same era.
Multi-domain Dialog State Tracking using Recurrent Neural Networks
Nikola Mrkšić, Diarmuid Ó Séaghdha, Blaise Thomson, Milica Gašić, Pei-Hao Su, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2015 · 2015
Cited alongside, same era.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, Yishay Mansour, et al. 1999a
Cited in the paper.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh. 1999b
Cited in the paper.
End-to-end lstm-based dialog control optimized with supervised and reinforcement learning
Jason D Williams and Geoffrey Zweig. 2016 · 2016
Later among the works it cites.
The Option-Critic Architecture
P.-L. Bacon, J. Harb, and D. Precup. 2017 · 2017
Closest in time.
Composite Task-Completion Dialogue System via Hierarchical Deep Reinforcement Learning
B. Peng, X. Li, L. Li, J. Gao, A. Celikyilmaz, S. Lee, and K.-F. Wong. 2017 · 2017
Closest in time.
Sample-efficient actor-critic reinforcement learning with supervised data for dialogue management
Pei-Hao Su, Paweł Budzianowski, Stefan Ultes, Milica Gašić, and Steve J. Young. 2017 · 2017
Closest in time.
Pydial: A multi-domain statistical dialogue system toolkit
Stefan Ultes, Lina M. Rojas-Barahona, Pei-Hao Su, David Vandyke, Dongho Kim, Iñigo Casanueva, Paweł Budzianowski, Nikola Mrkšić, Tsung-Hsien Wen, Milica Gašić, and Steve J. Young. 2017 · 2017
Closest in time.