Fetching the paper…
Reading the bibliography…
Conventional reinforcement learning (RL) needs an environment to collect fresh data, which is impractical when online interactions are costly.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N.; Ghandeharioun, A.; Shen, J. H.; Ferguson, C.; Lapedriza, A.; Jones, N.; Gu, S.; and Picard, R. 2019 · 1907
Earlier work this paper cites.
Behavior Regularized Offline Reinforcement Learning
Wu, Y.; Tucker, G.; and Nachum, O. 2019 · 1911
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J.; Kumar, A.; Nachum, O.; Tucker, G.; and Levine, S. 2020 · 2004
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 2005
Earlier work this paper cites.
MOReL : Model-Based Offline Reinforcement Learning
Kidambi, R.; Rajeswaran, A.; Netrapalli, P.; and Joachims, T. 2020 · 2005
Earlier work this paper cites.
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Levine, S.; Kumar, A.; Tucker, G.; and Fu, J. 2020 · 2005
Earlier work this paper cites.
MOPO: Model-based Offline Policy Optimization
Yu, T.; Thomas, G.; Yu, L.; Ermon, S.; Zou, J.; Levine, S.; Finn, C.; and Ma, T. 2020 · 2005
Earlier work this paper cites.
Accelerating Online Reinforcement Learning with Offline Datasets
Nair, A.; Dalal, M.; Gupta, A.; and Levine, S. 2020 · 2006
Earlier work this paper cites.
Supervised machine learning: A review of classification techniques
Kotsiantis, S. B.; Zaharakis, I.; Pintelas, P.; et al. 2007 · 2007
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S.; and Barto, A. G. 2011 · 2011
Earlier work this paper cites.
POPO: Pessimistic Offline Policy Optimization
He, Q.; and Hou, X. 2020 · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Cited alongside, same era.
Learning from Limited Demonstrations
Kim, B.; massoud Farahmand, A.; Pineau, J.; and Precup, D. 2013 · 2013
Cited alongside, same era.
Policy evaluation with temporal differences: A survey and comparison
Dann, C.; Neumann, G.; Peters, J.; et al. 2014 · 2014
Cited alongside, same era.
Offline policy evaluation across representations with applications to educational games
Mandel, T.; Liu, Y.-E.; Levine, S.; Brunskill, E.; and Popovic, Z. 2014 · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Rajeswaran, A.; Kumar, V.; Gupta, A.; Vezzani, G.; Schulman, J.; Todorov, E.; and Levine, S. 2018 · 2018
Later among the works it cites.
Off-Policy Deep Reinforcement Learning without Exploration
Fujimoto, S.; Meger, D.; and Precup, D. 2019 · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A.; Fu, J.; Soh, M.; Tucker, G.; and Levine, S. 2019 · 2019
Later among the works it cites.
Dexterous manipulation with deep reinforcement learning: Efficient, general, and low-cost
Zhu, H.; Gupta, A.; Rajeswaran, A.; Levine, S.; and Kumar, V. 2019 · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R.; Schuurmans, D.; and Norouzi, M. 2020 · 2020
Later among the works it cites.
Conservative Q-Learning for Offline Reinforcement Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trust region policy optimization
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; and Moritz, P. 2015 · 2015
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Vecerik, M.; Hester, T.; Scholz, J.; Wang, F.; Pietquin, O.; Piot, B.; Heess, N.; Rothörl, T.; Lampe, T.; and Riedmiller, M. 2017 · 2017
Cited alongside, same era.
Addressing Function Approximation Error in Actor-Critic Methods
Fujimoto, S.; van Hoof, H.; and Meger, D. 2018 · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Chen, X.; Wang, C.; Zhou, Z.; and Ross, K. W. 2021a
Cited in the paper.
Kumar, A.; Zhou, A.; Tucker, G.; and Levine, S. 2020 · 2020
Later among the works it cites.
A Minimalist Approach to Offline Reinforcement Learning
Fujimoto, S.; and Gu, S. S. 2021 · 2021
Later among the works it cites.
Offline-to-Online Reinforcement Learning via Balanced Replay and Pessimistic Q-Ensemble
Lee, S.; Seo, Y.; Lee, K.; Abbeel, P.; and Shin, J. 2021 · 2021
Later among the works it cites.
Deployment-efficient reinforcement learning via model-based offline optimization
Matsushima, T.; Furuta, H.; Matsuo, Y.; Nachum, O.; and Gu, S. 2021 · 2021
Later among the works it cites.
Offline rl policies should be trained to be adaptive
Ghosh, D.; Ajay, A.; Agrawal, P.; and Levine, S. 2022 · 2022
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Kostrikov, I.; Nair, A.; and Levine, S. 2022 · 2022
Later among the works it cites.