Fetching the paper…
Reading the bibliography…
Offline reinforcement learning seeks to utilize offline (observational) data to guide the learning of (causal) sequential decision making strategies.
Sequential analysis and optimal design
H. Chernoff · 1972
Earlier work this paper cites.
Feature-based methods for large scale dynamic programming
B. Van Roy · 1994
Earlier work this paper cites.
Neuro-dynamic programming: an overview
D. P. Bertsekas and J. N. Tsitsiklis · 1995
Earlier work this paper cites.
Stable function approximation in dynamic programming
G. J. Gordon · 1995
Earlier work this paper cites.
Stable fitted reinforcement learning
G. J. Gordon · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation (technical report lids-p-2322)
J. Tsitsiklis and B. Van Roy · 1996
Earlier work this paper cites.
An elementary introduction to modern convex geometry
K. Ball · 1997
Earlier work this paper cites.
Approximate solutions to markov decision processes
G. J. Gordon · 1999
Earlier work this paper cites.
Barycentric interpolators for continuous space and time reinforcement learning
R. Munos and A. W. Moore · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup · 2000
Earlier work this paper cites.
Kernel-based reinforcement learning
D. Ormoneit and Ś. Sen · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. M. Kakade · 2003
Earlier work this paper cites.
Least-squares policy iteration
M. G. Lagoudakis and R. Parr · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
R. Munos · 2003
Earlier work this paper cites.
Q ⋆ Q^{\star} approximation schemes for batch reinforcement learning: A theoretical comparison
T. Xie and N. Jiang · 2003
Earlier work this paper cites.
The sample complexity of exploration in the multi-armed bandit problem
S. Mannor and J. N. Tsitsiklis · 2004
Earlier work this paper cites.
Finite time bounds for sampling based fitted value iteration
C. Szepesvári and R. Munos · 2005
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
A. Antos, C. Szepesvári, and R. Munos · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
R. Munos and C. Szepesvári · 2008
Earlier work this paper cites.
Batch value-function approximation with only realizability
T. Xie and N. Jiang · 2008
Earlier work this paper cites.
Neural network learning: Theoretical foundations
M. Anthony and P. L. Bartlett · 2009
Earlier work this paper cites.
Doubly robust policy evaluation and learning
M. Dudík, J. Langford, and L. Li · 2011
Cited alongside, same era.
Towards minimax policies for online linear optimization with bandit feedback
S. Bubeck, N. Cesa-Bianchi, and S. M. Kakade · 2012
Cited alongside, same era.
Agnostic system identification for model-based reinforcement learning
S. Ross and D. Bagnell · 2012
Cited alongside, same era.
Offline policy evaluation across representations with applications to educational games
T. Mandel, Y.-E. Liu, S. Levine, E. Brunskill, and Z. Popovic · 2014
Cited alongside, same era.
Safe reinforcement learning
P. S. Thomas · 2014
Cited alongside, same era.
Toward minimax off-policy value estimation
L. Li, R. Munos, and C. Szepesvari · 2015
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Later among the works it cites.
N. Kallus and M. Uehara · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine · 2019
Later among the works it cites.
Safe policy improvement with baseline bootstrapping
R. Laroche, P. Trichelair, and R. Tachet des Combes · 2019
Later among the works it cites.
Understanding the curse of horizon in off-policy evaluation via conditional importance sampling
Y. Liu, P.-L. Bacon, and E. Brunskill · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
High-confidence off-policy evaluation
P. S. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Cited alongside, same era.
An introduction to matrix concentration inequalities
J. A. Tropp · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Cited alongside, same era.
Pac reinforcement learning with rich observations
A. Krishnamurthy, A. Agarwal, and J. Langford · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
P. Thomas and E. Brunskill · 2016
Cited alongside, same era.
Preventing undesirable behavior of intelligent machines
P. S. Thomas, B. Castro da Silva, A. G. Barto, S. Giguere, Y. Brun, and E. Brunskill · 2019
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
M. Uehara and N. Jiang · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
T. Xie, Y. Ma, and Y.-X. Wang · 2019
Later among the works it cites.
Deep inverse reinforcement learning for sepsis treatment
C. Yu, G. Ren, and J. Liu · 2019
Later among the works it cites.
Limiting extrapolation in linear approximate value iteration
A. Zanette, A. Lazaric, M. J. Kochenderfer, and E. Brunskill · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
R. Agarwal, D. Schuurmans, and M. Norouzi · 2020
Closest in time.
Is a good representation sufficient for sample efficient reinforcement learning?
S. S. Du, S. M. Kakade, R. Wang, and L. F. Yang · 2020
Closest in time.
Accountable off-policy evaluation with kernel bellman statistics
Y. Feng, T. Ren, Z. Tang, and Q. Liu · 2020
Closest in time.
Way off-policy batch deep reinforcement learning of human preferences in dialog, 2020
N. Jaques, A. Ghandeharioun, J. H. Shen, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard · 2020
Closest in time.
Minimax confidence interval for off-policy evaluation and policy optimization
N. Jiang and J. Huang · 2020
Closest in time.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
N. Kallus and M. Uehara · 2020
Closest in time.
Morel : Model-based offline reinforcement learning, 2020
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Closest in time.
Discor: Corrective feedback in reinforcement learning via distribution correction
A. Kumar, A. Gupta, and S. Levine · 2020
Closest in time.
On the sample complexity of reinforcement learning with policy space generalization
W. Mou, Z. Wen, and X. Chen · 2020
Closest in time.
On reward-free reinforcement learning with linear function approximation
R. Wang, S. S. Du, L. F. Yang, and R. Salakhutdinov · 2020
Closest in time.
Behavior regularized offline reinforcement learning, 2020
Y. Wu, G. Tucker, and O. Nachum · 2020
Closest in time.