Fetching the paper…
Reading the bibliography…
We study the regret of reinforcement learning from offline data generated by a fixed behavior policy in an infinite-horizon discounted Markov decision process (MDP).
Empirical processes: theory and applications
David Pollard · 1990
Earlier work this paper cites.
An upper bound on the loss from approximate optimal-value functions
Satinder P Singh and Richard C Yee · 1994
Earlier work this paper cites.
Smooth discrimination analysis
Enno Mammen and Alexandre B Tsybakov · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup, R. Sutton, and S Singh · 2000
Earlier work this paper cites.
Error bounds for approximate policy iteration
Rémi Munos · 2003
Earlier work this paper cites.
Least-squares policy iteration
Michail Lagoudakis and Ronald Parr · 2004
Earlier work this paper cites.
Optimal aggregation of classifiers in statistical learning
Alexander B Tsybakov · 2004
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Error bounds for approximate value iteration
Rémi Munos · 2005
Earlier work this paper cites.
Fast learning rates for plug-in classifiers
Jean-Yves Audibert and Alexandre B Tsybakov · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Finite-sample analysis of lstd
Alessandro Lazaric, Mohammad Ghavamzadeh, and Remi Munos · 2010
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J. Kappen · 2013
Cited alongside, same era.
A linear response bandit problem
Alexander Goldenshluger and Assaf Zeevi · 2013
Cited alongside, same era.
The multi-armed bandit problem with covariates
Vianney Perchet and Philippe Rigollet · 2013
Cited alongside, same era.
Approximate policy iteration schemes: A comparison
Bruno Scherrer · 2014
Cited alongside, same era.
An introduction to matrix concentration inequalities
Joel A Tropp · 2015
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
P. Thomas and E. Brunskill · 2016
Cited alongside, same era.
Online decision making with high-dimensional covariates
Hamsa Bastani and Mohsen Bayati · 2020
Later among the works it cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S Du, Sham M Kakade, Ruosong Wang, and Lin F Yang · 2020
Later among the works it cites.
Minimax-optimal off-policy evaluation with linear function approximation
Yaqi Duan, Zeyu Jia, and Mengdi Wang · 2020
Later among the works it cites.
A theoretical analysis of deep q-learning
Jianqing Fan, Zhaoran Wang, Yuchen Xie, and Zhuoran Yang · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
Provably good batch off-policy reinforcement learning without great exploration
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SBEED: Convergent reinforcement learning with nonlinear function approximation
Bo Dai, Albert Shaw, Lihong Li, Lin Xiao, Niao He, Zhen Liu, Jianshu Chen, and Le Song · 2018
Cited alongside, same era.
Breaking the curse of horizon: infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Cited alongside, same era.
Provably efficient q-learning with function approximation via distribution shift error checking oracle
Simon S Du, Yuping Luo, Ruosong Wang, and Hanrui Zhang · 2019
Cited alongside, same era.
Algaedice: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 2019
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Max Simchowitz and Kevin G Jamieson · 2019
Cited alongside, same era.
Later among the works it cites.
Performance guarantees for policy learning
Alex Luedtke and Antoine Chambaz · 2020
Later among the works it cites.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Tengyang Xie and Nan Jiang · 2020
Later among the works it cites.
What are the statistical limits of offline rl with linear function approximation?
Ruosong Wang, Dean Foster, and Sham M Kakade · 2021
Closest in time.
Q-learning with logarithmic regret
Kunhe Yang, Lin Yang, and Simon Du · 2021
Closest in time.
Near-optimal provable uniform convergence in offline policy evaluation for reinforcement learning
Ming Yin, Yu Bai, and Yu-Xiang Wang · 2021
Closest in time.
Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning
Nathan Kallus and Masatoshi Uehara · 2022
Closest in time.
Refined value-based offline rl under realizability and partial coverage
Masatoshi Uehara, Nathan Kallus, Jason D Lee, and Wen Sun · 2023
Closest in time.