Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) enables the agent to effectively learn from logged data, which significantly extends the applicability of RL algorithms in real-world scenarios where exploration can be expensive or unsafe.
Dynamics-aware unsupervised discovery of skills
Sharma, A.; Gu, S.; Levine, S.; Kumar, V.; and Hausman, K. 2019 · 1907
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Gupta, A.; Kumar, V.; Lynch, C.; Levine, S.; and Hausman, K. 2019 · 1910
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B.; Kumar, A.; Zhang, G.; and Levine, S. 2019 · 1910
Earlier work this paper cites.
The MAXQ Method for Hierarchical Reinforcement Learning
Dietterich, T. G.; et al. 1998 · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S.; Precup, D.; and Singh, S. 1999 · 1999
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J.; Kumar, A.; Nachum, O.; Tucker, G.; and Levine, S. 2020 · 2004
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S.; Kumar, A.; Tucker, G.; and Fu, J. 2020 · 2005
Earlier work this paper cites.
Accelerating online reinforcement learning with offline datasets
Nair, A.; Dalal, M.; Gupta, A.; and Levine, S. 2020 · 2006
Earlier work this paper cites.
Argenson, A.; and Dulac-Arnold, G. 2020 · 2008
Earlier work this paper cites.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
Ajay, A.; Kumar, A.; Agrawal, P.; Levine, S.; and Nachum, O. 2020 · 2010
Earlier work this paper cites.
Accelerating reinforcement learning with learned skill priors
Pertsch, K.; Lee, Y.; and Lim, J. J. 2020 · 2010
Earlier work this paper cites.
Cog: Connecting new skills to past experience with offline reinforcement learning
Singh, A.; Yu, A.; Yang, J.; Zhang, J.; Kumar, A.; and Levine, S. 2020b · 2010
Earlier work this paper cites.
Parrot: Data-driven behavioral priors for reinforcement learning
Singh, A.; Liu, H.; Zhou, G.; Yu, A.; Rhinehart, N.; and Levine, S. 2020a · 2011
Earlier work this paper cites.
Nice: Non-linear independent components estimation
Dinh, L.; Krueger, D.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2015 · 2015
Cited alongside, same era.
Batch learning from logged bandit feedback through counterfactual risk minimization
Swaminathan, A.; and Joachims, T. 2015 · 2015
Cited alongside, same era.
Density estimation using real nvp
Dinh, L.; Sohl-Dickstein, J.; and Bengio, S. 2016 · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D.; Narasimhan, K.; Saeedi, A.; and Tenenbaum, J. 2016 · 2016
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S.; Shammah, S.; and Shashua, A. 2016 · 2016
Mastering complex control in moba games with deep reinforcement learning
Ye, D.; Liu, Z.; Sun, M.; Shi, B.; Zhao, P.; Wu, H.; Yu, H.; Yang, S.; Wu, X.; Guo, Q.; et al. 2020 · 2020
Later among the works it cites.
Offline rl without off-policy evaluation
Brandfonbrener, D.; Whitney, W.; Ranganath, R.; and Bruna, J. 2021 · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S.; and Gu, S. S. 2021 · 2021
Later among the works it cites.
Emaq: Expected-max q-learning operator for simple yet effective offline and online rl
Ghasemipour, S. K. S.; Schuurmans, D.; and Gu, S. S. 2021 · 2021
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Jin, Y.; Yang, Z.; and Wang, Z. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distributional reinforcement learning with quantile regression
Dabney, W.; Rowland, M.; Bellemare, M.; and Munos, R. 2018 · 2018
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B.; Gupta, A.; Ibarz, J.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Near-optimal representation learning for hierarchical reinforcement learning
Nachum, O.; Gu, S.; Lee, H.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Yang, L.; and Wang, M. 2019 · 2019
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C.; Yang, Z.; Wang, Z.; and Jordan, M. I. 2020 · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
Kidambi, R.; Rajeswaran, A.; Netrapalli, P.; and Joachims, T. 2020 · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A.; Zhou, A.; Tucker, G.; and Levine, S. 2020 · 2020
Cited alongside, same era.
Kostrikov, I.; Nair, A.; and Levine, S. 2021 · 2021
Later among the works it cites.
Offline Reinforcement Learning with Value-based Episodic Memory
Ma, X.; Yang, Y.; Hu, H.; Liu, Q.; Yang, J.; Zhang, C.; Zhao, Q.; and Liang, B. 2021 · 2021
Later among the works it cites.
Conservative offline distributional reinforcement learning
Ma, Y.; Jayaraman, D.; and Bastani, O. 2021 · 2021
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M.; and Sun, W. 2021 · 2021
Later among the works it cites.
Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning
Yang, Y.; Ma, X.; Chenghao, L.; Zheng, Z.; Zhang, Q.; Huang, G.; Yang, J.; and Zhao, Q. 2021 · 2021
Later among the works it cites.
Combo: Conservative offline model-based policy optimization
Yu, T.; Kumar, A.; Rafailov, R.; Rajeswaran, A.; Levine, S.; and Finn, C. 2021 · 2021
Later among the works it cites.
On the Role of Discount Factor in Offline Reinforcement Learning
Hu, H.; Yang, Y.; Zhao, Q.; and Zhang, C. 2022 · 2022
Closest in time.
Skill-based Meta-Reinforcement Learning
Nam, T.; Sun, S.-H.; Pertsch, K.; Hwang, S. J.; and Lim, J. J. 2022 · 2022
Closest in time.