Fetching the paper…
Reading the bibliography…
An important aspect of developing conversational agents is to give a bot the ability to improve through communicating with humans and to learn from the mistakes that it makes.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Instructive feedback: Review of parameters and effects
Margaret G Werts, Mark Wolery, Ariane Holcombe, and David L Gast · 1995
Earlier work this paper cites.
Learning dialogue strategies within the markov decision process framework
Esther Levin, Roberto Pieraccini, and Wieland Eckert · 1997
Earlier work this paper cites.
A stochastic model of human-machine interaction for learning dialog strategies
Esther Levin, Roberto Pieraccini, and Wieland Eckert · 2000
Earlier work this paper cites.
Empirical evaluation of a reinforcement learning spoken dialogue system
Satinder Singh, Michael Kearns, Diane J Litman, Marilyn A Walker, et al · 2000
Earlier work this paper cites.
An application of reinforcement learning to dialogue strategy selection in a spoken dialogue system for email
Marilyn A. Walker · 2000
Earlier work this paper cites.
Optimizing dialogue management with reinforcement learning: Experiments with the njfun system
Satinder Singh, Diane Litman, Michael Kearns, and Marilyn Walker · 2002
Earlier work this paper cites.
A trainable generator for recommendations in multimodal dialog
Marilyn A Walker, Rashmi Prasad, and Amanda Stent · 2003
Earlier work this paper cites.
A survey of statistical user simulation techniques for reinforcement-learning of dialogue management strategies
Jost Schatzmann, Karl Weilhammer, Matt Stuttle, and Steve Young · 2006
Earlier work this paper cites.
Are we there yet? research in commercial spoken dialog systems
Roberto Pieraccini, David Suendermann, Krishna Dayanidhi, and Jackson Liscombe · 2009
Earlier work this paper cites.
The hidden information state model: A practical framework for pomdp-based spoken dialogue management
Steve Young, Milica Gašić, Simon Keizer, François Mairesse, Jost Schatzmann, Blaise Thomson, and Kai Yu · 2010
Cited alongside, same era.
Interactional feedback and the impact of attitude and motivation on noticing l2 form
Mohammad Amin Bassiri · 2011
Cited alongside, same era.
Counterfactual reasoning and learning systems: The example of computational advertising
Leon Bottou, Jonas Peters, Denis X. Quiñonero-Candela, Joaquin amd Charles, D. Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Cited alongside, same era.
Pomdp-based dialogue manager adaptation to extended domains
Milica Gašic, Catherine Breslin, Matthew Henderson, Dongho Kim, Martin Szummer, Blaise Thomson, Pirros Tsiakoulis, and Steve Young · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Later among the works it cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2015
Later among the works it cites.
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al · 2015
Later among the works it cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart van Merriënboer, Armand Joulin, and Tomas Mikolov · 2015
Later among the works it cites.
Key-value memory networks for directly reading documents
Alexander Miller, Adam Fisch, Jesse Dodge, Amir-Hossein Karimi, Antoine Bordes, and Jason Weston · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pomdp-based statistical spoken dialog systems: A review
Steve Young, Milica Gašić, Blaise Thomson, and Jason D Williams · 2013
Cited alongside, same era.
Incremental on-line adaptation of pomdp-based dialogue managers to extended domains
Milica Gašic, Dongho Kim, Pirros Tsiakoulis, Catherine Breslin, Matthew Henderson, Martin Szummer, Blaise Thomson, and Steve Young · 2014
Cited alongside, same era.
Large-scale simple question answering with memory networks
Antoine Bordes, Nicolas Usunier, Sumit Chopra, and Jason Weston · 2015
Cited alongside, same era.
Evaluating prerequisite qualities for learning end-to-end dialog systems
Jesse Dodge, Andreea Gane, Xiang Zhang, Antoine Bordes, Sumit Chopra, Alexander Miller, Arthur Szlam, and Jason Weston · 2015
Cited alongside, same era.
Closest in time.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Closest in time.
Continuously learning neural dialogue management
Pei-Hao Su, Milica Gasic, Nikola Mrksic, Lina Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young · 2016
Closest in time.
Dialog-based language learning
Jason Weston · 2016
Closest in time.
Reinforcement learning neural turing machines
Wojciech Zaremba and Ilya Sutskever · 2016
Closest in time.