Fetching the paper…
Reading the bibliography…
Offline Reinforcement Learning (RL) aims to learn a near-optimal policy from a fixed dataset of transitions collected by another policy.
Linear programming and sequential decisions
Manne, A. S · 1909
Earlier work this paper cites.
Dynamic programming
Bellman, R · 1956
Earlier work this paper cites.
Dynamic programming
Bellman, R · 1966
Earlier work this paper cites.
The extragradient method for finding saddle points and other problems
Korpelevich, G · 1976
Earlier work this paper cites.
Constrained Optimization and Lagrange Multiplier Methods
Bertsekas, D. P · 1982
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
Nemirovski, A. and Yudin, D · 1983
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
Markov Chains and Stochastic Stability
Meyn, S. and Tweedie, R · 1996
Earlier work this paper cites.
Stochastic approximation with two time scales
Borkar, V. S · 1997
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Prediction, Learning, and Games
Cesa-Bianchi, N. and Lugosi, G · 2006
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Earlier work this paper cites.
Q-learning and pontryagin’s minimum principle
Mehta, P. G. and Meyn, S. P · 2009
Cited alongside, same era.
Learning bounds for importance weighting
Cortes, C., Mansour, Y., and Mohri, M · 2010
Cited alongside, same era.
Optimization, learning, and games with predictable sequences
Rakhlin, A. and Sridharan, K · 2013
Cited alongside, same era.
An online primal-dual method for discounted markov decision processes
Wang, M. and Chen, Y · 2016
Cited alongside, same era.
A unified view of entropy-regularized Markov decision processes
Neu, G., Jonsson, A., and Gómez, V · 2017
Cited alongside, same era.
Scalable bilinear learning using state and action features
Chen, Y., Li, L., and Wang, M · 2018
Cited alongside, same era.
Minimax weight and q-function learning for off-policy evaluation
Uehara, M., Huang, J., and Jiang, N · 2020
Later among the works it cites.
Logistic q-learning
Bas-Serrano, J., Curi, S., Krause, A., and Neu, G · 2021
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Jin, Y., Yang, Z., and Wang, Z · 2021
Later among the works it cites.
Corruption-robust exploration in episodic reinforcement learning
Lykouris, T., Simchowitz, M., Slivkins, A., and Sun, W · 2021
Later among the works it cites.
On the optimality of batch policy optimization algorithms
Xiao, C., Wu, Y., Mei, J., Dai, B., Lattimore, T., Li, L., Szepesvári, C., and Schuurmans, D · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P., and Agarwal, A · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A modern introduction to online learning
Orabona, F · 2019
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Yang, L. and Wang, M · 2019
Cited alongside, same era.
Faster saddle-point optimization for solving large-scale markov decision processes
Bas-Serrano, J. and Neu, G · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
Provably good batch off-policy reinforcement learning without great exploration
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E · 2020
Cited alongside, same era.
Provable benefits of actor-critic methods for offline reinforcement learning
Zanette, A., Wainwright, M. J., and Brunskill, E · 2021
Later among the works it cites.
Adversarially trained actor critic for offline reinforcement learning
Cheng, C.-A., Xie, T., Jiang, N., and Agarwal, A · 2022
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S · 2022
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M. and Sun, W · 2022
Later among the works it cites.
Corruption-robust offline reinforcement learning
Zhang, X., Chen, Y., Zhu, X., and Sun, W · 2022
Later among the works it cites.
Online learning with off-policy feedback
Gabbianelli, G., Neu, G., and Papini, M · 2023
Closest in time.
Efficient global planning in large mdps via stochastic primal-dual optimization
Neu, G. and Okolo, N · 2023
Closest in time.