Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) enables effective learning from previously collected data without exploration, which shows great promise in real-world applications when exploration is expensive or even infeasible.
Empirical Processes in M-estimation , volume 6
Geer, S. A., van de Geer, S., and Williams, D · 2000
Earlier work this paper cites.
Biasing approximate dynamic programming with a lower discount factor
Petrik, M. and Scherrer, B · 2008
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
The dependence of effective planning horizon on model accuracy
Jiang, N., Kulesza, A., Singh, S., and Lewis, R · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Batch learning from logged bandit feedback through counterfactual risk minimization
Swaminathan, A. and Joachims, T · 2015
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S., Shammah, S., and Shashua, A · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Improving offline value-function approximations for pomdps by reducing discount factors
Chen, Y.-C., Kochenderfer, M. J., and Spaan, M. T · 2018
Earlier work this paper cites.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H. P., and Silver, D · 2018
Earlier work this paper cites.
Hyperbolic discounting and learning over multiple horizons
Fedus, W., Gelada, C., Bengio, Y., Bellemare, M. G., and Larochelle, H · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Earlier work this paper cites.
Separating value functions across time-scales
Romoff, J., Henderson, P., Touati, A., Brunskill, E., Pineau, J., and Ollivier, Y · 2019
Earlier work this paper cites.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Wainwright, M. J · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Yang, L. and Wang, M · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Cited alongside, same era.
Discount factor as a regularizer in reinforcement learning
Amit, R., Meir, R., and Ciosek, K · 2020
Cited alongside, same era.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T · 2020
Later among the works it cites.
A deeper look at discounting mismatch in actor-critic algorithms
Zhang, S., Laroche, R., van Seijen, H., Whiteson, S., and Combes, R. T. d · 2020
Later among the works it cites.
Uncertainty-based offline reinforcement learning with diversified q-ensemble
An, G., Moon, S., Kim, J.-H., and Song, H. O · 2021
Later among the works it cites.
Offline rl without off-policy evaluation
Brandfonbrener, D., Whitney, W. F., Ranganath, R., and Bruna, J · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Argenson, A. and Dulac-Arnold, G · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Hong, M., Wai, H.-T., Wang, Z., and Yang, Z · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
Fujimoto, S. and Gu, S. S · 2021
Later among the works it cites.
Emaq: Expected-max q-learning operator for simple yet effective offline and online rl
Ghasemipour, S. K. S., Schuurmans, D., and Gu, S. S · 2021
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Jin, Y., Yang, Z., and Wang, Z · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2021
Later among the works it cites.
Offline reinforcement learning with value-based episodic memory
Ma, X., Yang, Y., Hu, H., Liu, Q., Yang, J., Zhang, C., Zhao, Q., and Liang, B · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S · 2021
Later among the works it cites.
Comparison and unification of three regularization methods in batch reinforcement learning
Rathnam, S., Murphy, S. A., and Doshi-Velez, F · 2021
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M. and Sun, W · 2021
Later among the works it cites.
Uncertainty weighted actor-critic for offline reinforcement learning
Wu, Y., Zhai, S., Srivastava, N., Susskind, J., Zhang, J., Salakhutdinov, R., and Goh, H · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P., and Agarwal, A · 2021
Later among the works it cites.
Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning
Yang, Y., Ma, X., Li, C., Zheng, Z., Zhang, Q., Huang, G., Yang, J., and Zhao, Q · 2021
Later among the works it cites.
Combo: Conservative offline model-based policy optimization
Yu, T., Kumar, A., Rafailov, R., Rajeswaran, A., Levine, S., and Finn, C · 2021
Later among the works it cites.