Fetching the paper…
Reading the bibliography…
We study the problem of safe offline reinforcement learning (RL), the goal is to learn a policy that maximizes long-term reward while satisfying safety constraints given only offline data, without further interaction with the environment.
Behavior Regularized Offline Reinforcement Learning
Wu, Y.; Tucker, G.; and Nachum, O. 2019 · 1911
Earlier work this paper cites.
Q-learning
Watkins, C. J.; and Dayan, P. 1992 · 1992
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S.; Barto, A. G.; et al. 1998 · 1998
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Altman, E. 1999 · 1999
Earlier work this paper cites.
Learning to Walk in the Real World with Minimal Human Effort
Ha, S.; Xu, P.; Tan, Z.; Levine, S.; and Tan, J. 2020 · 2002
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S.; Kumar, A.; Tucker, G.; and Fu, J. 2020 · 2005
Earlier work this paper cites.
Conservative Q-Learning for Offline Reinforcement Learning
Kumar, A.; Zhou, A.; Tucker, G.; and Levine, S. 2020 · 2006
Earlier work this paper cites.
Fitted Q-iteration in continuous action-space MDPs
Antos, A.; Szepesvári, C.; and Munos, R. 2008 · 2008
Earlier work this paper cites.
Batch reinforcement learning
Lange, S.; Gabel, T.; and Riedmiller, M. 2012 · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. 2013 · 2013
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcıa, J.; and Fernández, F. 2015 · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; and Moritz, P. 2015 · 2015
Earlier work this paper cites.
End-to-End Training of Deep Visuomotor Policies
Levine, S.; Finn, C.; Darrell, T.; and Abbeel, P. 2016 · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2016 · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J.; Held, D.; Tamar, A.; and Abbeel, P. 2017 · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Chow, Y.; Ghavamzadeh, M.; Janson, L.; and Pavone, M. 2017 · 2017
Cited alongside, same era.
beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework
Higgins, I.; Matthey, L.; Pal, A.; Burgess, C. P.; Glorot, X.; Botvinick, M. M.; Mohamed, S.; and Lerchner, A. 2017 · 2017
Cited alongside, same era.
Batch Policy Learning under Constraints
Le, H.; Voloshin, C.; and Yue, Y. 2019 · 2019
Later among the works it cites.
Reinforcement learning with convex constraints
Miryoosefi, S.; Brantley, K.; Daume III, H.; Dudik, M.; and Schapire, R. E. 2019 · 2019
Later among the works it cites.
Likelihood ratios for out-of-distribution detection
Ren, J.; Liu, P. J.; Fertig, E.; Snoek, J.; Poplin, R.; Depristo, M.; Dillon, J.; and Lakshminarayanan, B. 2019 · 2019
Later among the works it cites.
MOReL: Model-Based Offline Reinforcement Learning
Kidambi, R.; Rajeswaran, A.; Netrapalli, P.; and Joachims, T. 2020 · 2020
Later among the works it cites.
Energy-based Out-of-distribution Detection
Liu, W.; Wang, X.; Owens, J.; and Li, Y. 2020 · 2020
Later among the works it cites.
Constrained markov decision processes via backward value functions
Satija, H.; Amortila, P.; and Pineau, J. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D.; Schrittwieser, J.; Simonyan, K.; Antonoglou, I.; Huang, A.; Guez, A.; Hubert, T.; Baker, L.; Lai, M.; Bolton, A.; et al. 2017 · 2017
Cited alongside, same era.
Safe exploration in continuous action spaces
Dalal, G.; Dvijotham, K.; Vecerik, M.; Hester, T.; Paduraru, C.; and Tassa, Y. 2018 · 2018
Cited alongside, same era.
Addressing Function Approximation Error in Actor-Critic Methods
Fujimoto, S.; Hoof, H.; and Meger, D. 2018 · 2018
Cited alongside, same era.
Accelerated primal-dual policy optimization for safe reinforcement learning
Liang, Q.; Que, F.; and Modiano, E. 2018 · 2018
Cited alongside, same era.
Reward Constrained Policy Optimization
Tessler, C.; Mankowitz, D. J.; and Mannor, S. 2018 · 2018
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A.; Fu, J.; Soh, M.; Tucker, G.; and Levine, S. 2019 · 2019
Cited alongside, same era.
MOPO: Model-based Offline Policy Optimization
Yu, T.; Thomas, G.; Yu, L.; Ermon, S.; Zou, J.; Levine, S.; Finn, C.; and Ma, T. 2020 · 2020
Later among the works it cites.
Model-Based Offline Planning
Argenson, A.; and Dulac-Arnold, G. 2021 · 2021
Closest in time.
Zhan, X.; Xu, H.; Zhang, Y.; Huo, Y.; Zhu, X.; Yin, H.; and Zheng, Y. 2021 · 2021
Closest in time.
Model-Based Offline Planning with Trajectory Pruning
Zhan, X.; Zhu, X.; and Xu, H. 2021 · 2021
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S.; Meger, D.; and Precup, D. 2019 · 2062
Closest in time.