Fetching the paper…
Reading the bibliography…
Offline reinforcement learning proposes to learn policies from large collected datasets without interacting with the physical environment.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Kernel-based reinforcement learning
D. Ormoneit and Ś. Sen · 2002
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
D. Ernst, P. Geurts, and L. Wehenkel · 2005
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
M. Riedmiller · 2005
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2017
Earlier work this paper cites.
Adversarially robust policy learning: Active construction of physically-plausible perturbations
A. Mandlekar, Y. Zhu, A. Garg, L. Fei-Fei, and S. Savarese · 2017
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2017
Earlier work this paper cites.
ROBOTURK: A Crowdsourcing Platform for Robotic Skill Learning through Imitation
A. Mandlekar, Y. Zhu, A. Garg, J. Booher, M. Spero, A. Tung, J. Gao, J. Emmons, A. Gupta, E. Orbay, S. Savarese, and L. Fei-Fei · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, G. Tucker, and S. Levine · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and O. Nachum · 2019
Cited alongside, same era.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
N. Jaques, A. Ghandeharioun, J. H. Shen, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard · 2019
Cited alongside, same era.
A framework for data-driven robotics
S. Cabi, S. G. Colmenarejo, A. Novikov, K. Konyushkova, S. Reed, R. Jeong, K. Zołna, Y. Aytar, D. Budden, M. Vecerik, et al · 2019
Cited alongside, same era.
Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulation dataset through human reasoning and dexterity
A. Mandlekar, J. Booher, M. Spero, A. Tung, A. Gupta, Y. Zhu, A. Garg, S. Savarese, and L. Fei-Fei · 2019
Cited alongside, same era.
Curl: Contrastive unsupervised representations for reinforcement learning
M. Laskin, A. Srinivas, and P. Abbeel · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
I. Kostrikov, D. Yarats, and R. Fergus · 2020
Later among the works it cites.
Reinforcement learning with augmented data
M. Laskin, K. Lee, A. Stooke, L. Pinto, P. Abbeel, and A. Srinivas · 2020
Later among the works it cites.
Experience replay with likelihood-free importance weights
S. Sinha, J. Song, A. Garg, and S. Ermon · 2020
Later among the works it cites.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn · 2019
Cited alongside, same era.
Improving sample efficiency in model-free reinforcement learning from images
D. Yarats, A. Zhang, I. Kostrikov, B. Amos, J. Pineau, and R. Fergus · 2019
Cited alongside, same era.
Mixmatch: A holistic approach to semi-supervised learning
D. Berthelot, N. Carlini, I. Goodfellow, N. Papernot, A. Oliver, and C. Raffel · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y. Yoo · 2019
Cited alongside, same era.
robosuite: A modular simulation framework and benchmark for robot learning
Y. Zhu, J. Wong, A. Mandlekar, and R. Martín-Martín · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Cited alongside, same era.
Rl unplugged: Benchmarks for offline reinforcement learning
C. Gulcehre, Z. Wang, A. Novikov, T. L. Paine, S. G. Colmenarejo, K. Zolna, R. Agarwal, J. Merel, D. Mankowitz, C. Paduraru, et al · 2020
Cited alongside, same era.
A. Ajay, A. Kumar, P. Agrawal, S. Levine, and O. Nachum · 2020
Later among the works it cites.
The importance of pessimism in fixed-dataset policy optimization
J. Buckman, C. Gelada, and M. G. Bellemare · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Later among the works it cites.
Cog: Connecting new skills to past experience with offline reinforcement learning
A. Singh, A. Yu, J. Yang, J. Zhang, A. Kumar, and S. Levine · 2020
Later among the works it cites.
J. Jin, D. Graves, C. Haigh, J. Luo, and M. Jagersand · 2020
Later among the works it cites.
D2rl: Deep dense architectures in reinforcement learning
S. Sinha, H. Bharadhwaj, A. Srinivas, and A. Garg · 2020
Later among the works it cites.
Data-efficient reinforcement learning with momentum predictive representations
M. Schwarzer, A. Anand, R. Goel, R. D. Hjelm, A. Courville, and P. Bachman · 2020
Later among the works it cites.