Fetching the paper…
Reading the bibliography…
Recent advance in deep offline reinforcement learning (RL) has made it possible to train strong robotic agents from offline datasets.
Learning attractor landscapes for learning motor primitives
A. J. Ijspeert, J. Nakanishi, and S. Schaal · 2003
Earlier work this paper cites.
Sampling techniques
W. G. Cochran · 2007
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by penalized convex risk minimization
X. Nguyen, M. J. Wainwright, and M. I. Jordan · 2008
Earlier work this paper cites.
Double q-learning
H. V. Hasselt · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Learning from limited demonstrations
B. Kim, A.-m. Farahmand, J. Pineau, and D. Precup · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Learning structured output representation using deep conditional generative models
K. Sohn, H. Lee, and X. Yan · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
J. Garcıa and F. Fernández · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
H. v. Hasselt, A. Guez, and D. Silver · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Earlier work this paper cites.
Importance weighted autoencoders
Y. Burda, R. Grosse, and R. Salakhutdinov · 2016
Earlier work this paper cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. Riedmiller · 2017
Earlier work this paper cites.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
O. Anschel, N. Baram, and N. Shimkin · 2017
Earlier work this paper cites.
Ucb exploration via q-ensembles
R. Y. Chen, S. Sidor, P. Abbeel, and J. Schulman · 2017
Cited alongside, same era.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Hoof, and D. Meger · 2018
Cited alongside, same era.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Cog: Connecting new skills to past experience with offline reinforcement learning
A. Singh, A. Yu, J. Yang, J. Zhang, A. Kumar, and S. Levine · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Hyperparameter selection for offline reinforcement learning
T. L. Paine, C. Paduraru, A. Michi, C. Gulcehre, K. Zolna, A. Novikov, Z. Wang, and N. de Freitas · 2020
Later among the works it cites.
Emaq: Expected-max q-learning operator for simple yet effective offline and online rl
S. K. S. Ghasemipour, D. Schuurmans, and S. S. Gu · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Ebert, C. Finn, S. Dasari, A. Xie, A. Lee, and S. Levine · 2018
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and O. Nachum · 2019
Cited alongside, same era.
Dexterous manipulation with deep reinforcement learning: Efficient, general, and low-cost
H. Zhu, A. Gupta, A. Rajeswaran, S. Levine, and V. Kumar · 2019
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Cited alongside, same era.
Later among the works it cites.
Accelerating reinforcement learning with learned skill priors
K. Pertsch, Y. Lee, and J. J. Lim · 2020
Later among the works it cites.
Maxmin q-learning: Controlling the estimation bias of q-learning
Q. Lan, Y. Pan, A. Fyshe, and M. White · 2020
Later among the works it cites.
Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning
K. Lee, M. Laskin, A. Srinivas, and P. Abbeel · 2020
Later among the works it cites.
Deployment-efficient reinforcement learning via model-based offline optimization
T. Matsushima, H. Furuta, Y. Matsuo, O. Nachum, and S. Gu · 2020
Later among the works it cites.
Experience replay with likelihood-free importance weights, 2021
S. Sinha, J. Song, A. Garg, and S. Ermon · 2021
Closest in time.
Awac: Accelerating online reinforcement learning with offline datasets
A. Nair, M. Dalal, A. Gupta, and S. Levine · 2021
Closest in time.
{OPAL}: Offline primitive discovery for accelerating offline reinforcement learning
A. Ajay, A. Kumar, P. Agrawal, S. Levine, and O. Nachum · 2021
Closest in time.
Parrot: Data-driven behavioral priors for reinforcement learning
A. Singh, H. Liu, G. Zhou, A. Yu, N. Rhinehart, and S. Levine · 2021
Closest in time.
Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning
K. Lee, M. Laskin, A. Srinivas, and P. Abbeel · 2021
Closest in time.
A minimalist approach to offline reinforcement learning
S. Fujimoto and S. Gu · 2021
Closest in time.