Fetching the paper…
Reading the bibliography…
KL-regularized reinforcement learning from expert demonstrations has proved successful in improving the sample efficiency of deep reinforcement learning algorithms, allowing them to be applied to challenging physical real-world tasks.
The theory of Tikhonov regularization for Fredholm equations
CW Groetsch · 1984
Earlier work this paper cites.
A framework for behavioural cloning
Michael Bain and Claude Sammut · 1995
Earlier work this paper cites.
Behavioural cloning: phenomena, results and problems
Ivan Bratko, Tanja Urbancic, and Claude Sammut · 1995
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore · 1996
Earlier work this paper cites.
Learning from demonstration
Stefan Schaal et al · 1997
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y. Ng and Stuart J. Russell · 2000
Earlier work this paper cites.
Supervised actor-critic reinforcement learning
Michael T Rosenstein, Andrew G Barto, Jennie Si, Andy Barto, and Warren Powell · 2004
Earlier work this paper cites.
Natural actor-critic
Jan Peters, Sethu Vijayakumar, and Stefan Schaal · 2005
Earlier work this paper cites.
A unifying view of sparse approximate Gaussian process regression
Joaquin Quiñonero Candela and Carl Edward Rasmussen · 2005
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
Carl Edward Rasmussen and Christopher K. I. Williams · 2005
Earlier work this paper cites.
A hilbert space embedding for distributions
Alex Smola, Arthur Gretton, Le Song, and Bernhard Schölkopf · 2007
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov · 2007
Earlier work this paper cites.
Double Q-learning
Hado V. Hasselt · 2010
Earlier work this paper cites.
Relative entropy inverse reinforcement learning
Abdeslam Boularias, Jens Kober, and Jan Peters · 2011
Earlier work this paper cites.
Off-policy actor-critic
Thomas Degris, Martha White, and Richard S. Sutton · 2012
Cited alongside, same era.
Robot learning from demonstration by constructing skill trees
George Konidaris, Scott Kuindersma, Roderic Grupen, and Andrew Barto · 2012
Cited alongside, same era.
Real-world reinforcement learning for autonomous humanoid robot docking
Nicolás Navarro-Guerrero, Cornelius Weber, Pascal Schroeter, and Stefan Wermter · 2012
Cited alongside, same era.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Cited alongside, same era.
Machine learning: A Probabilistic Perspective
Soft actor-critic algorithms and applications, 2019
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine · 2019
Later among the works it cites.
Stabilizing off-policy Q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Later among the works it cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning, 2019
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Later among the works it cites.
Exact Gaussian processes on a million data points
Ke Wang, Geoff Pleiss, Jacob Gardner, Stephen Tyree, Kilian Q Weinberger, and Andrew Gordon Wilson · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kevin P. Murphy · 2013
Cited alongside, same era.
Reinforcement learning from demonstrations through shaping
Tim Brys, Anna Harutyunyan, Halit Bener Suay, Sonia Chernova, Matthew E Taylor, and Ann Nowé · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Cited alongside, same era.
Equivalence between policy gradients and soft Q-learning
John Schulman, Xi Chen, and Pieter Abbeel · 2017
Cited alongside, same era.
Reinforcement learning from imperfect demonstrations
Yang Gao, Huazhe Xu, Ji Lin, Fisher Yu, Sergey Levine, and Trevor Darrell · 2018
Cited alongside, same era.
Reinforcement learning and control as probabilistic inference: Tutorial and review, 2018
Sergey Levine · 2018
Cited alongside, same era.
Radial bayesian neural networks: Beyond discrete support in large-scale bayesian deep learning
Sebastian Farquhar, Michael A. Osborne, and Yarin Gal · 2020
Later among the works it cites.
MOReL: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Later among the works it cites.
Accelerating online reinforcement learning with offline datasets, 2020
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2020
Later among the works it cites.
Learning to score behaviors for guided policy optimization
Aldo Pacchiano, Jack Parker-Holder, Yunhao Tang, Krzysztof Choromanski, Anna Choromanska, and Michael Jordan · 2020
Later among the works it cites.
Keep doing what worked: Behavior modelling priors for offline reinforcement learning
Noah Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin Riedmiller · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Augmented world models facilitate zero-shot dynamics generalization from a single offline environment
Philip J Ball, Cong Lu, Jack Parker-Holder, and Stephen Roberts · 2021
Later among the works it cites.
Behavioral priors and dynamics models: Improving performance and domain transfer in offline RL, 2021
Catherine Cang, Aravind Rajeswaran, Pieter Abbeel, and Michael Laskin · 2021
Later among the works it cites.
Revisiting design choices in model-based offline reinforcement learning, 2021
Cong Lu, Philip J. Ball, Jack Parker-Holder, Michael A. Osborne, and Stephen J. Roberts · 2021
Later among the works it cites.
Improving deterministic uncertainty estimation in deep learning for classification and regression, 2021
Joost van Amersfoort, Lewis Smith, Andrew Jesson, Oscar Key, and Yarin Gal · 2021
Later among the works it cites.
COMBO: Conservative offline model-based policy optimization, 2021
Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn · 2021
Later among the works it cites.