Fetching the paper…
Reading the bibliography…
Most deep reinforcement learning (RL) systems are not able to learn effectively from off-policy data, especially if they cannot explore online in the environment.
Stochastic optimal control
Robert F Stengel · 1986
Earlier work this paper cites.
Learning to achieve goals
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Laughter
Robert R Provine · 1996
Earlier work this paper cites.
Functions of humor in the conversations of men and women
Jennifer Hay · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Where to look: a study of human-robot engagement
Candace L Sidner, Cory D Kidd, Christopher Lee, and Neal Lesh · 2004
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mülling, and Yasemin Altun · 2010
Earlier work this paper cites.
Active listening in peer interviews: The influence of message paraphrasing on perceptions of listening skill
Harry Weger Jr, Gina R Castle, and Melissa C Emmett · 2010
Earlier work this paper cites.
Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs
Cristian Danescu-Niculescu-Mizil and Lillian Lee · 2011
Earlier work this paper cites.
On-line policy optimisation of spoken dialogue systems via live interaction with human subjects
Milica Gašić, Filip Jurčíček, Blaise Thomson, Kai Yu, and Steve Young · 2011
Earlier work this paper cites.
Language style matching predicts relationship initiation and stability
Molly E Ireland, Richard B Slatcher, Paul W Eastwick, Lauren E Scissors, Eli J Finkel, and James W Pennebaker · 2011
Earlier work this paper cites.
Listening competence in initial interactions i: Distinguishing between what listening is and what listeners do
Graham D Bodie, Kellie St. Cyr, Michelle Pence, Michael Rold, and James Honeycutt · 2012
Earlier work this paper cites.
Off-policy actor-critic
Thomas Degris, Martha White, and Richard S Sutton · 2012
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Hilbert J Kappen, Vicenç Gómez, and Manfred Opper · 2012
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
The role of “active listening” in informal helping conversations: Impact on perceptions of listener helpfulness, sensitivity, and supportiveness and discloser emotional improvement
Graham D Bodie, Andrea J Vickery, Kaitlin Cannava, and Susanne M Jones · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Policy networks with two-stage training for dialogue systems
Mehdi Fatemi, Layla El Asri, Hannes Schulz, Jing He, and Kaheer Suleman · 2016
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
Microsoft deletes ’teen girl’ ai after it became a hitler-loving sex robot within 24 hours
Helena Horton · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
A hierarchical latent variable encoder-decoder model for generating dialogues
Iulian Vlad Serban, Alessandro Sordoni, Ryan Lowe, Laurent Charlin, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2017
Later among the works it cites.
Sample-efficient actor-critic reinforcement learning with supervised data for dialogue management
Pei-Hao Su, Paweł Budzianowski, Stefan Ultes, Milica Gasic, and Steve Young · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Later among the works it cites.
Efficient exploration through bayesian deep q-networks
Kamyar Azizzadenesheli, Emma Brunskill, and Animashree Anandkumar · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dialogue learning with human-in-the-loop
Jiwei Li, Alexander H Miller, Sumit Chopra, Marc’Aurelio Ranzato, and Jason Weston · 2016
Cited alongside, same era.
Deep reinforcement learning for dialogue generation
Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Building end-to-end dialogue systems using generative hierarchical neural network models
Iulian V Serban, Alessandro Sordoni, Yoshua Bengio, Aaron Courville, and Joelle Pineau · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Dialogue generation: From imitation learning to inverse reinforcement learning
Ziming Li, Julia Kiseleva, and Maarten de Rijke · 2018
Later among the works it cites.
Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems
Bing Liu, Gokhan Tür, Dilek Hakkani-Tür, Pararth Shah, and Larry Heck · 2018
Later among the works it cites.
A hierarchical latent structure for variational conversation modeling
Yookoon Park, Jaemin Cho, and Gunhee Kim · 2018
Later among the works it cites.
Bootstrapping a neural conversational agent with dialogue self-play, crowdsourcing and on-line reinforcement learning
Pararth Shah, Dilek Hakkani-Tur, Bing Liu, and Gokhan Tur · 2018
Later among the works it cites.
Sentiment adaptive end-to-end dialog systems
Weiyan Shi and Zhou Yu · 2018
Later among the works it cites.
The design and implementation of xiaoice, an empathetic social chatbot
Li Zhou, Jianfeng Gao, Di Li, and Heung-Yeung Shum · 2018
Later among the works it cites.
Crossnorm: Normalization for off-policy td reinforcement learning
Aditya Bhatt, Max Argus, Artemij Amiranashvili, and Thomas Brox · 2019
Closest in time.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Carles Gelada and Marc G Bellemare · 2019
Closest in time.
Approximating interactive human evaluation with self-play for open-domain dialog systems
Asma Ghandeharioun, Judy Shen, Natasha Jaques, Craig Ferguson, Noah Jones, Agata Lapedriza, and Rosalind Picard · 2019
Closest in time.
Learning from dialogue after deployment: Feed yourself, chatbot!
Braden Hancock, Antoine Bordes, Pierre-Emmanuel Mazare, and Jason Weston · 2019
Closest in time.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Closest in time.
Happybot: Generating empathetic dialogue responses by improving user experience look-ahead
Jamin Shin, Peng Xu, Andrea Madotto, and Pascale Fung · 2019
Closest in time.