Fetching the paper…
Reading the bibliography…
A compelling use case of offline reinforcement learning (RL) is to obtain a policy initialization from existing datasets followed by fast online fine-tuning with limited interaction.
Learning from demonstration
S. Schaal · 1996
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Accelerating online reinforcement learning with offline datasets
A. Nair, M. Dalal, A. Gupta, and S. Levine · 2006
Earlier work this paper cites.
Accelerating online reinforcement learning with offline datasets
A. Nair, M. Dalal, A. Gupta, and S. Levine · 2006
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2017
Earlier work this paper cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. Riedmiller · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Deep q-learning from demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, et al · 2018
Earlier work this paper cites.
Policy optimization with demonstrations
B. Kang, Z. Jie, and J. Feng · 2018
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2018
Earlier work this paper cites.
Reinforcement and imitation learning for diverse visuomotor skills
Y. Zhu, Z. Wang, J. Merel, A. Rusu, T. Erez, S. Cabi, S. Tunyasuvunakool, J. Kramár, R. Hadsell, N. de Freitas, et al · 2018
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2019
Earlier work this paper cites.
Algaedice: Policy gradient from arbitrary experience
O. Nachum, B. Dai, I. Kostrikov, Y. Chow, L. Li, and D. Schuurmans · 2019
Earlier work this paper cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Earlier work this paper cites.
Dexterous manipulation with deep reinforcement learning: Efficient, general, and low-cost
H. Zhu, A. Gupta, A. Rajeswaran, S. Levine, and V. Kumar · 2019
Earlier work this paper cites.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
A. Ajay, A. Kumar, P. Agrawal, S. Levine, and O. Nachum · 2020
Earlier work this paper cites.
The importance of pessimism in fixed-dataset policy optimization
J. Buckman, C. Gelada, and M. G. Bellemare · 2020
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Batch reinforcement learning through continuation method
Y. Guo, S. Feng, N. Le Roux, E. Chi, H. Lee, and M. Chen · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Bandit algorithms
T. Lattimore and C. Szepesvári · 2020
Cited alongside, same era.
Jaxcql: a simple implementation of sac and cql in jax
X. Geng · 2022
Later among the works it cites.
Unpacking reward shaping: Understanding the benefits of reward engineering on sample complexity
A. Gupta, A. Pacchiano, Y. Zhai, S. M. Kakade, and S. Levine · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Later among the works it cites.
Pre-Training for Robots: Offline RL Enables Learning New Tasks from a Handful of Trials
A. Kumar, A. Singh, F. Ebert, Y. Yang, C. Finn, and S. Levine · 2022
Later among the works it cites.
Pre-training for robots: Offline rl enables learning new tasks from a handful of trials
A. Kumar, A. Singh, F. Ebert, Y. Yang, C. Finn, and S. Levine · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Cited alongside, same era.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
N. Y. Siegel, J. T. Springenberg, F. Berkenkamp, A. Abdolmaleki, M. Neunert, T. Lampe, R. Hafner, and M. Riedmiller · 2020
Cited alongside, same era.
Cog: Connecting new skills to past experience with offline reinforcement learning
A. Singh, A. Yu, J. Yang, J. Zhang, A. Kumar, and S. Levine · 2020
Cited alongside, same era.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
T. Xie and N. Jiang · 2020
Cited alongside, same era.
Towards playing full moba games with deep reinforcement learning
D. Ye, G. Chen, W. Zhang, S. Chen, B. Yuan, B. Liu, J. Chen, Z. Liu, F. Qiu, H. Yu, Y. Yin, B. Shi, L. Wang, T. Shi, Q. Fu, W. Yang, L. Huang, and W. Liu · 2020
Cited alongside, same era.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Cited alongside, same era.
Randomized ensembled double q-learning: Learning fast without a model
X. Chen, C. Wang, Z. Zhou, and K. W. Ross · 2021
Cited alongside, same era.
Later among the works it cites.
Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble
S. Lee, Y. Seo, K. Lee, P. Abbeel, and J. Shin · 2022
Later among the works it cites.
Understanding the complexity gains of single-task rl with a curriculum
Q. Li, Y. Zhai, Y. Ma, and S. Levine · 2022
Later among the works it cites.
Mildly conservative q-learning for offline reinforcement learning
J. Lyu, X. Ma, X. Li, and Z. Lu · 2022
Later among the works it cites.
Fine-tuning offline policies with optimistic action selection
M. S. Mark, A. Ghadirzadeh, X. Chen, and C. Finn · 2022
Later among the works it cites.
Leveraging offline data in online reinforcement learning
A. Wagenmaker and A. Pacchiano · 2022
Later among the works it cites.
Supported policy optimization for offline reinforcement learning
J. Wu, H. Wu, Z. Qiu, J. Wang, and M. Long · 2022
Later among the works it cites.
Computational benefits of intermediate rewards for goal-reaching policy learning
Y. Zhai, C. Baek, Z. Zhou, J. Jiao, and Y. Ma · 2022
Later among the works it cites.
Online decision transformer
Q. Zheng, A. Zhang, and A. Grover · 2022
Later among the works it cites.
Efficient online reinforcement learning with offline data
P. J. Ball, L. Smith, I. Kostrikov, and S. Levine · 2023
Closest in time.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
P. D’Oro, M. Schwarzer, E. Nikishin, P.-L. Bacon, M. G. Bellemare, and A. Courville · 2023
Closest in time.
Efficient deep reinforcement learning requires regulating overfitting
Q. Li, A. Kumar, I. Kostrikov, and S. Levine · 2023
Closest in time.
Hybrid RL: Using both offline and online data can make RL efficient
Y. Song, Y. Zhou, A. Sekhari, D. Bagnell, A. Krishnamurthy, and W. Sun · 2023
Closest in time.
The in-sample softmax for offline reinforcement learning
C. Xiao, H. Wang, Y. Pan, A. White, and M. White · 2023
Closest in time.