Fetching the paper…
Reading the bibliography…
Most prior approaches to offline reinforcement learning (RL) utilize \textit{behavior regularization}, typically augmenting existing off-policy actor critic algorithms with a penalty measuring divergence between the policy and the offline data.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B.; Kumar, A.; Zhang, G.; and Levine, S. 2019 · 1910
Earlier work this paper cites.
Behavior Regularized Offline Reinforcement Learning
Wu, Y.; Tucker, G.; and Nachum, O. 2019 · 1911
Earlier work this paper cites.
Algaedice: Policy gradient from arbitrary experience
Nachum, O.; Dai, B.; Kostrikov, I.; Chow, Y.; Li, L.; and Schuurmans, D. 2019b · 1912
Earlier work this paper cites.
Nonlinear programming
Bertsekas, D. P. 1997 · 1997
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S.; Barto, A. G.; et al. 1998 · 1998
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S.; and Langford, J. 2002 · 2002
Earlier work this paper cites.
BRPO: Batch Residual Policy Optimization
Sohn, S.; Chow, Y.; Ooi, J.; Nachum, O.; Lee, H.; Chi, E.; and Boutilier, C. 2020 · 2002
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J.; Kumar, A.; Nachum, O.; Tucker, G.; and Levine, S. 2020 · 2004
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S.; Kumar, A.; Tucker, G.; and Fu, J. 2020 · 2005
Earlier work this paper cites.
Accelerating online reinforcement learning with offline datasets
Nair, A.; Dalal, M.; Gupta, A.; and Levine, S. 2020 · 2006
Earlier work this paper cites.
Stable dual dynamic programming
Wang, T.; Bowling, M.; Schuurmans, D.; and Lizotte, D. J. 2008 · 2008
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
Nguyen, X.; Wainwright, M. J.; and Jordan, M. I. 2010 · 2010
Earlier work this paper cites.
Off-policy actor-critic
Degris, T.; White, M.; and Sutton, R. S. 2012 · 2012
Earlier work this paper cites.
A kernel two-sample test
Gretton, A.; Borgwardt, K. M.; Rasch, M. J.; Schölkopf, B.; and Smola, A. 2012 · 2012
Earlier work this paper cites.
Batch reinforcement learning
Lange, S.; Gabel, T.; and Riedmiller, M. 2012 · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. 2013 · 2013
Earlier work this paper cites.
Safe policy iteration
Pirotta, M.; Restelli, M.; Pecorino, A.; and Calandriello, D. 2013 · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Trust region policy optimization
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; and Moritz, P. 2015 · 2015
Earlier work this paper cites.
End-to-End Training of Deep Visuomotor Policies
Levine, S.; Finn, C.; Darrell, T.; and Abbeel, P. 2016 · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2016 · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J.; Held, D.; Tamar, A.; and Abbeel, P. 2017 · 2017
Cited alongside, same era.
Learning from conditional distributions via dual embeddings
Dai, B.; He, N.; Pan, Y.; Boots, B.; and Song, L. 2017 · 2017
Cited alongside, same era.
Consistent on-line off-policy evaluation
Hallak, A.; and Mannor, S. 2017 · 2017
Cited alongside, same era.
Markov chains and mixing times , volume 107
Levin, D. A.; and Peres, Y. 2017 · 2017
Cited alongside, same era.
Kernel Mean Embedding of Distributions: A Review and Beyond
Muandet, K.; Fukumizu, K.; Sriperumbudur, B.; and Schölkopf, B. 2017 · 2017
Cited alongside, same era.
DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction
Kumar, A.; Gupta, A.; and Levine, S. 2020 · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Kumar, A.; Zhou, A.; Tucker, G.; and Levine, S. 2020 · 2020
Later among the works it cites.
Batch reinforcement learning with hyperparameter gradients
Lee, B.; Lee, J.; Vrancx, P.; Kim, D.; and Kim, K.-E. 2020 · 2020
Later among the works it cites.
Black-box Off-policy Estimation for Infinite-Horizon Reinforcement Learning
Mousavi, A.; Li, L.; Liu, Q.; and Zhou, D. 2020 · 2020
Later among the works it cites.
Constrained markov decision processes via backward value functions
Satija, H.; Amortila, P.; and Pineau, J. 2020 · 2020
Later among the works it cites.
Truly proximal policy optimization
Wang, Y.; He, H.; and Tan, X. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rajeswaran, A.; Kumar, V.; Gupta, A.; Vezzani, G.; Schulman, J.; Todorov, E.; and Levine, S. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Reinforcement learning based recommender system using biclustering technique
Choi, S.; Ha, H.; Hwang, U.; Kim, C.; Ha, J.-W.; and Yoon, S. 2018 · 2018
Cited alongside, same era.
Addressing Function Approximation Error in Actor-Critic Methods
Fujimoto, S.; Hoof, H.; and Meger, D. 2018 · 2018
Cited alongside, same era.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q.; Li, L.; Tang, Z.; and Zhou, D. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Critic Regularized Regression
Wang, Z.; Novikov, A.; Zolna, K.; Merel, J. S.; Springenberg, J. T.; Reed, S. E.; Shahriari, B.; Siegel, N.; Gulcehre, C.; Heess, N.; et al. 2020 · 2020
Later among the works it cites.
BRAC+: Going Deeper with Behavior Regularized Offline Reinforcement Learning
Zhang, C.; Kuppannagari, S. R.; and Prasanna, V. 2020 · 2020
Later among the works it cites.
Latent Action Space for Offline Reinforcement Learning
Zhou, W.; Bajracharya, S.; and Held, D. 2020 · 2020
Later among the works it cites.
Offline RL Without Off-Policy Evaluation
Brandfonbrener, D.; Whitney, W. F.; Ranganath, R.; and Bruna, J. 2021 · 2021
Closest in time.
Offline Reinforcement Learning with Pseudometric Learning
Dadashi, R.; Rezaeifar, S.; Vieillard, N.; Hussenot, L.; Pietquin, O.; and Geist, M. 2021 · 2021
Closest in time.
A Minimalist Approach to Offline Reinforcement Learning
Fujimoto, S.; and Gu, S. S. 2021 · 2021
Closest in time.
Emaq: Expected-max q-learning operator for simple yet effective offline and online rl
Ghasemipour, S. K. S.; Schuurmans, D.; and Gu, S. S. 2021 · 2021
Closest in time.
Offline reinforcement learning with fisher divergence critic regularization
Kostrikov, I.; Fergus, R.; Tompson, J.; and Nachum, O. 2021 · 2021
Closest in time.
OptiDICE: Offline Policy Optimization via Stationary Distribution Correction Estimation
Lee, J.; Jeon, W.; Lee, B.-J.; Pineau, J.; and Kim, K.-E. 2021 · 2021
Closest in time.
Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning
Wu, Y.; Zhai, S.; Srivastava, N.; Susskind, J.; Zhang, J.; Salakhutdinov, R.; and Goh, H. 2021 · 2021
Closest in time.
Constraints Penalized Q-Learning for Safe Offline Reinforcement Learning
Xu, H.; Zhan, X.; and Zhu, X. 2021 · 2021
Closest in time.
Zhan, X.; Xu, H.; Zhang, Y.; Huo, Y.; Zhu, X.; Yin, H.; and Zheng, Y. 2021 · 2021
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S.; Meger, D.; and Precup, D. 2019 · 2062
Closest in time.