Fetching the paper…
Reading the bibliography…
We make progress in a long-standing problem of batch reinforcement learning (RL): learning $Q^\star$ from an exploratory and polynomial-sized dataset, using a realizable and otherwise arbitrary function class.
Approximations of dynamic programs, I
Whitt, W · 1978
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Stable function approximation in dynamic programming
Gordon, G. J · 1995
Earlier work this paper cites.
A generalized reinforcement-learning model: Convergence and applications
Littman, M. L. and Szepesvári, C · 1996
Earlier work this paper cites.
Eligibility Traces for Off-Policy Policy Evaluation
Precup, D., Sutton, R. S., and Singh, S. P · 2000
Earlier work this paper cites.
Approximate equivalence of Markov decision processes
Even-Dar, E. and Mansour, Y · 2003
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R · 2003
Earlier work this paper cites.
Approximate homomorphisms: A framework for nonexact minimization in Markov decision processes
Ravindran, B. and Barto, A · 2004
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Li, L., Walsh, T. J., and Littman, M. L · 2006
Earlier work this paper cites.
Performance bounds in l_p-norm for approximate value iteration
Munos, R · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., Szepesvári, C., and Munos, R · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Earlier work this paper cites.
Model-based reinforcement learning with state aggregation
Paduraru, C., Kaplow, R., Precup, D., and Pineau, J · 2008
Earlier work this paper cites.
Error Propagation for Approximate Policy and Value Iteration
Farahmand, A.-m., Szepesvári, C., and Munos, R · 2010
Earlier work this paper cites.
E0 370 Statistical Learning Theory: Covering Numbers, Pseudo-Dimension, and Fat-Shattering Dimension
Agarwal, S · 2011
Cited alongside, same era.
Model selection in reinforcement learning
Farahmand, A.-m. and Szepesvári, C · 2011
Cited alongside, same era.
Combinatorial methods in density estimation
Devroye, L. and Lugosi, G · 2012
Cited alongside, same era.
Model selection in markovian processes
Hallak, A., Di-Castro, D., and Mannor, S · 2013
Cited alongside, same era.
Policy iteration based on stochastic factorization
Barreto, A. d. M. S., Pineau, J., and Precup, D · 2014
Cited alongside, same era.
Offline policy evaluation across representations with applications to educational games
Mandel, T., Liu, Y.-E., Levine, S., Brunskill, E., and Popovic, Z · 2014
Cited alongside, same era.
On value functions and the agent-environment boundary
Jiang, N · 2019
Later among the works it cites.
Off-policy policy gradient with state distribution correction
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E · 2019
Later among the works it cites.
Model-based RL in Contextual Decision Processes: PAC bounds and Exponential Improvements over Model-free Approaches
Sun, W., Jiang, N., Krishnamurthy, A., Agarwal, A., and Langford, J · 2019
Later among the works it cites.
A variant of the wang-foster-kakade lower bound for the discounted setting
Amortila, P., Jiang, N., and Xie, T · 2020
Closest in time.
Accountable off-policy evaluation with kernel bellman statistics
Feng, Y., Ren, T., Tang, Z., and Liu, Q · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient abstraction selection in reinforcement learning
van Seijen, H., Whiteson, S., and Kester, L · 2014
Cited alongside, same era.
Abstraction Selection in Model-based Reinforcement Learning
Jiang, N., Kulesza, A., and Singh, S · 2015
Cited alongside, same era.
Doubly Robust Off-policy Value Evaluation for Reinforcement Learning
Jiang, N. and Li, L · 2016
Cited alongside, same era.
Contextual decision processes with low Bellman rank are PAC-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E · 2017
Cited alongside, same era.
CS 598: Notes on State Abstractions
Jiang, N · 2018
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., and Zhou, D · 2018
Cited alongside, same era.
Closest in time.
Minimax value interval for off-policy evaluation and policy optimization
Jiang, N. and Huang, J · 2020
Closest in time.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 2020
Closest in time.
Provably good batch reinforcement learning without great exploration
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E · 2020
Closest in time.
Hyperparameter selection for offline reinforcement learning
Paine, T. L., Paduraru, C., Michi, A., Gulcehre, C., Zolna, K., Novikov, A., Wang, Z., and de Freitas, N · 2020
Closest in time.
Minimax weight and q-function learning for off-policy evaluation
Uehara, M., Huang, J., and Jiang, N · 2020
Closest in time.
What are the statistical limits of offline rl with linear function approximation?
Wang, R., Foster, D. P., and Kakade, S. M · 2020
Closest in time.
Q ⋆ Q^{\star} Approximation Schemes for Batch Reinforcement Learning: A Theoretical Comparison
Xie, T. and Jiang, N · 2020
Closest in time.
Zanette, A · 2020
Closest in time.
Chen, L., Scherrer, B., and Bartlett, P. L · 2021
Closest in time.