Fetching the paper…
Reading the bibliography…
Learning from demonstrations is a popular tool for accelerating and reducing the exploration requirements of reinforcement learning.
Mind in Society: The Development of Higher Psychological Processes
L. S. Vygotsky · 1978
Earlier work this paper cites.
Self-Improving Reactive Agents Based on Reinforcement Learning, Planning and Teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
The Zone of Proximal Development in Vygotsky’s Analysis of Learning and Instruction
Seth Chaiklin · 2003
Earlier work this paper cites.
Curriculum Learning
Yoshua Bengio, Jerome Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
Stephane Ross, Geoffrey J Gordon, and J Andrew Bagnell · 2011
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2014
Earlier work this paper cites.
Boosted bellman residual minimization handling expert demonstrations
Bilal Piot, Matthieu Geist, and Olivier Pietquin · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Human-Level Control Through Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Continuous Control With Deep Reinforcement Learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Earlier work this paper cites.
Policy Distillation
Andrei A. Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2016
Cited alongside, same era.
Prioritized Experience Replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Cited alongside, same era.
Deep Reinforcement Learning With Double Q-Learning
Hado van Hasselt, Arthur Quez, and David Silver · 2016
Cited alongside, same era.
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2017
Cited alongside, same era.
Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
Matej Vecerik, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothorl, Thomas Lampe, and Martin Riedmiller · 2017
Cited alongside, same era.
Observe and Look Further: Achieving Consistent Performance on Atari
Tobias Pohlen, Bilal Piot, Todd Hester, Mohammad Gheshlaghi Azar, Dan Horgan, David Budden, Gabriel Barth-Maron, Hado van Hasselt, John Quan, Mel Večerík, Matteo Hessel, Rémi Munos, and Olivier Pietquin · 2018
Later among the works it cites.
Learning Montezuma’s Revenge From a Single Demonstration
Tim Salimans and Richard Chen · 2018
Later among the works it cites.
Kickstarting Deep Reinforcement Learning
Simon Schmitt, Jonathan J. Hudson, Augustin Zidek, Simon Osindero, Carl Doersch, Wojciech M. Czarnecki, Joel Z. Leibo, Heinrich Kuttler, Andrew Zisserman, Karen Simonyan, and S. M. Ali Eslami · 2018
Later among the works it cites.
Introduction to Reinforcement Learning
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
A Practical Approach to Insertion with Variable Socket Position Using Deep Reinforcement Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A Deeper Look at Experience Replay
Shangtong Zhang and Richard S. Sutton · 2017
Cited alongside, same era.
Born Again Neural Networks
Tommaso Furlanello, Zachary C. Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Cited alongside, same era.
Rainbow: Combining Improvements in Deep Reinforcement Learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Cited alongside, same era.
Deep Q-Learning From Demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Gabriel Dulac-Arnold, Ian Osband, John Agapiou, Joel Z. Leibo, and Audrunas Gruslys · 2018
Cited alongside, same era.
Distributed Prioritized Experience Replay
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado van Hasselt, and David Silver · 2018
Cited alongside, same era.
Overcoming Exploration in Reinforcement Learning with Demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Cited alongside, same era.
Mel Vecerik, Oleg Sushkov, David Barker, Thomas Rothorl, Todd Hester, and Jon Scholz · 2018
Later among the works it cites.
Striving for Simplicity in Off-policy Deep Reinforcement Learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2019
Closest in time.
Exploration by Random Network Distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2019
Closest in time.
Diagnosing Bottlenecks in Deep Q-learning Algorithms
Justin Fu, Aviral Kumar, Matthew Soh, and Sergey Levine · 2019
Closest in time.
Off-Policy Deep Reinforcement Learning without Exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Closest in time.
Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine · 2019
Closest in time.
Making Efficient Use of Demonstrations to Solve Hard Exploration Problems
Tom Le Paine, Caglar Gulcehre, Bobak Shahriari, Misha Denil, Matt Hoffman, Hubert Soyer nd Richard Tanburn, Steven Kapturowski, Neil Rabinowitz, Duncan Williams, Gabriel Barth-Maron, Ziyu Wang, Nando de Freitas, and Worlds Team · 2019
Closest in time.