Fetching the paper…
Reading the bibliography…
Pretraining with expert demonstrations have been found useful in speeding up the training process of deep reinforcement learning algorithms since less online simulation data is required.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
Richard S. Sutton, David Mcallester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Ng and Stuart Russell · 2000
Earlier work this paper cites.
Approximately Optimal Approximate Reinforcement Learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
Umar Syed and Robert E Schapire · 2008
Earlier work this paper cites.
Apprenticeship learning using linear programming
Umar Syed, Michael Bowling, and Robert E Schapire · 2008
Earlier work this paper cites.
Natural actor-critic algorithms
Shalabh Bhatnagar, Richard Sutton, Mohammad Ghavamzadeh, and Mark Lee · 2009
Earlier work this paper cites.
Off-Policy Actor-Critic.pdf
Thomas Degris, Martha White, and Richard S Sutton · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Boosted bellman residual minimization handling expert demonstrations
Bilal Piot, Matthieu Geist, and Olivier Pietquin · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Tim Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Model-free imitation learning with policy optimization
Jonathan Ho, Jayesh Gupta, and Stefano Ermon · 2016
Later among the works it cites.
Reinforcement Learning with Few Expert Demonstrations
Aravind S Lakshminarayanan, Sherjil Ozair, and Yoshua Bengio · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Safe and Efficient Off-Policy Reinforcement Learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G. Bellemare · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Trust Region Policy Optimization
John Schulman, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Generative Adversarial Imitation Learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas · 2016
Later among the works it cites.
Pre-training neural networks with human demonstrations for deep reinforcement learning
Gabriel V. de la Cruz, Jr., Yunshu Du, and Matthew E. Taylor · 2017
Later among the works it cites.
Q-Prop : Sample-Efficient Policy Gradient with An Off -Policy Critic
Shixiang Gu, Timothy Lillicrap, Zoubin Ghahramani, Richard E Turner, and Sergey Levine · 2017
Later among the works it cites.
Learning from demonstrations for real world reinforcement learning
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Andrew Sendonaris, Gabriel Dulac-Arnold, Ian Osband, John Agapiou, et al · 2017
Later among the works it cites.