Fetching the paper…
Reading the bibliography…
Developing agents to engage in complex goal-oriented dialogues is challenging partly because the main learning signals are very sparse in long conversations.
Finding structure in reinforcement learning
Sebastian Thrun and Anton Schwartz. 1994 · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Topic detection and tracking pilot study final report
James Allan, Jaime G Carbonell, George Doddington, Jonathan Yamron, and Yiming Yang. 1998 · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto. 1998 · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh. 1999 · 1999
Earlier work this paper cites.
An activity based approach to pragmatics
Jens Allwood. 2000 · 2000
Earlier work this paper cites.
A stochastic model of human-machine interaction for learning dialog strategies
Esther Levin, Roberto Pieraccini, and Wieland Eckert. 2000 · 2000
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
Amy McGovern and Andrew G Barto. 2001 · 2001
Earlier work this paper cites.
PolicyBlocks: An algorithm for creating useful macro-actions in reinforcement learning
Marc Pickett and Andrew G. Barto. 2002 · 2002
Earlier work this paper cites.
Learning options in reinforcement learning
Martin Stolle and Doina Precup. 2002 · 2002
Earlier work this paper cites.
Dynamic abstraction in reinforcement learning via clustering
Shie Mannor, Ishai Menache, Amit Hoze, and Uri Klein. 2004 · 2004
Earlier work this paper cites.
Chinese word segmentation and named entity recognition: A pragmatic approach
Jianfeng Gao, Mu Li, Andi Wu, and Chang-Ning Huang. 2005 · 2005
Earlier work this paper cites.
Identifying useful subgoals in reinforcement learning by local graph partitioning
Özgür Şimşek, Alicia P. Wolfe, and Andrew G. Barto. 2005 · 2005
Earlier work this paper cites.
Agenda-based user simulation for bootstrapping a pomdp dialogue system
Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve Young. 2007 · 2007
Earlier work this paper cites.
Topic segmentation with shared topic detection and alignment of multiple documents
Bingjun Sun, Prasenjit Mitra, C Lee Giles, John Yen, and Hongyuan Zha. 2007 · 2007
Earlier work this paper cites.
Rethinking Language, Mind, and World Dialogically
Per Linell. 2009 · 2009
Cited alongside, same era.
Evaluation of a hierarchical reinforcement learning spoken dialogue system
Heriberto Cuayáhuitl, Steve Renals, Oliver Lemon, and Hiroshi Shimodaira. 2010 · 2010
Cited alongside, same era.
Subgoal discovery in reinforcement learning using local graph clustering
Negin Entezari, Mohammad Ebrahim Shiri, and Parham Moradi. 2011 · 2011
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton. 2012 · 2012
Cited alongside, same era.
On the bottleneck concept for options discovery: Theoretical underpinnings and extension in continuous state spaces
Pierre-Luc Bacon. 2013 · 2013
Cited alongside, same era.
PAC-inspired option discovery in lifelong reinforcement learning
Tiancheng Zhao and Maxine Eskenazi. 2016 · 2016
Later among the works it cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup. 2017 · 2017
Later among the works it cites.
Sub-domain modelling for dialogue management with hierarchical reinforcement learning
Pawel Budzianowski, Stefan Ultes, Pei-Hao Su, Nikola Mrksic, Tsung-Hsien Wen, Inigo Casanueva, Lina Rojas-Barahona, and Milica Gasic. 2017 · 2017
Later among the works it cites.
Frames: A corpus for adding memory to goal-oriented dialogue systems
Layla El Asri, Hannes Schulz, Shikhar Sharma, Jeremie Zumer, Justin Harris, Emery Fine, Rahul Mehrotra, and Kaheer Suleman. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emma Brunskill and Lihong Li. 2014 · 2014
Cited alongside, same era.
Distributed dialogue policies for multi-domain statistical dialogue management
Milica Gašić, Dongho Kim, Pirros Tsiakoulis, and Steve Young. 2015 · 2015
Cited alongside, same era.
Policy committee for adaptation in multi-domain spoken dialogue systems
Milica Gašić, Nikola Mrkšić, Pei hao Su, David Vandyke, Tsung-Hsien Wen, and Steve J. Young. 2015 · 2015
Cited alongside, same era.
Machine learning for dialog state tracking: A review
Matthew Henderson. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. 2015 · 2015
Cited alongside, same era.
Deep reinforcement learning for multi-domain dialogue systems
Heriberto Cuayáhuitl, Seunghak Yu, Ashley Williamson, and Jacob Carse. 2016 · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum. 2016 · 2016
Cited alongside, same era.
Xuijun Li, Yun-Nung Chen, Lihong Li, and Jianfeng Gao. 2017 · 2017
Later among the works it cites.
Iterative policy learning in end-to-end trainable task-oriented neural dialog models
Bing Liu and Ian Lane. 2017 · 2017
Later among the works it cites.
End-to-end optimization of task-oriented dialogue model with deep reinforcement learning
Bing Liu, Gokhan Tur, Dilek Hakkani-Tur, Pararth Shah, and Larry Heck. 2017 · 2017
Later among the works it cites.
A Laplacian framework for option discovery in reinforcement learning
Marlos C. Machado, Marc G. Bellemare, and Michael H. Bowling. 2017 · 2017
Later among the works it cites.
Neural belief tracker: Data-driven dialogue state tracking
Nikola Mrkšić, Diarmuid Ó Séaghdha, Tsung-Hsien Wen, Blaise Thomson, and Steve J. Young. 2017 · 2017
Later among the works it cites.
Composite task-completion dialogue policy learning via hierarchical deep reinforcement learning
Baolin Peng, Xiujun Li, Lihong Li, Jianfeng Gao, Asli Celikyilmaz, Sungjin Lee, and Kam-Fai Wong. 2017b · 2017
Later among the works it cites.
FeUdal Networks for hierarchical reinforcement learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu. 2017 · 2017
Later among the works it cites.
Sequence modeling via segmentations
Chong Wang, Yining Wang, Po-Sen Huang, Abdelrahman Mohamed, Dengyong Zhou, and Li Deng. 2017 · 2017
Later among the works it cites.
Hybrid code networks: Practical and efficient end-to-end dialog control with supervised and reinforcement learning
Jason D Williams, Kavosh Asadi, and Geoffrey Zweig. 2017 · 2017
Later among the works it cites.
Integrating planning for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Kam-Fai Wong. 2018 · 2018
Closest in time.