Fetching the paper…
Reading the bibliography…
We study the offline reinforcement learning (offline RL) problem, where the goal is to learn a reward-maximizing policy in an unknown Markov Decision Process (MDP) using the data coming from a policy $\mu$.
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
Herman Chernoff et al · 1952
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
A gentle introduction to concentration inequalities
Karthik Sridharan · 2002
Earlier work this paper cites.
An adaptation theory for nonparametric confidence intervals
T Tony Cai and Mark G Low · 2004
Earlier work this paper cites.
Finite time bounds for sampling based fitted value iteration
Csaba Szepesvári and Rémi Munos · 2005
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Gaussian processes for sample efficient reinforcement learning with rmax-like exploration
Tobias Jung and Peter Stone · 2010
Earlier work this paper cites.
Freedman’s inequality for matrix martingales
Joel Tropp et al · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Efficient exploration and value function generalization in deterministic systems
Zheng Wen and Benjamin Van Roy · 2013
Earlier work this paper cites.
How hard is my mdp?” the distribution-norm to the rescue”
Odalric-Ambrym Maillard, Timothy A Mann, and Shie Mannor · 2014
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Earlier work this paper cites.
PAC reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Nan Jiang and Alekh Agarwal · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Problem dependent reinforcement learning bounds which can identify bandit structure in mdps
Andrea Zanette and Emma Brunskill · 2018
Cited alongside, same era.
Provably efficient q-learning with low switching cost
Yu Bai, Tengyang Xie, Nan Jiang, and Yu-Xiang Wang · 2019
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Cited alongside, same era.
Batch policy learning under constraints
Hoang Le, Cameron Voloshin, and Yisong Yue · 2019
Cited alongside, same era.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Cited alongside, same era.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Tengyang Xie and Nan Jiang · 2020
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Ming Yin and Yu-Xiang Wang · 2020
Later among the works it cites.
The importance of pessimism in fixed-dataset policy optimization
Jacob Buckman, Carles Gelada, and Marc G Bellemare · 2021
Closest in time.
Mitigating covariate shift in imitation learning via offline data without great coverage
Jonathan D Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi, and Wen Sun · 2021
Closest in time.
Benchmarks for deep off-policy evaluation
Justin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker, Ziyu Wang, Alexander Novikov, Mengjiao Yang, Michael R Zhang, Yutian Chen, Aviral Kumar, et al · 2021
Closest in time.
Reinforcement learning as one big sequence modeling problem
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Cited alongside, same era.
Model-based reinforcement learning with a generative model is minimax optimal
Alekh Agarwal, Sham Kakade, and Lin F Yang · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Cited alongside, same era.
Reinforcement learning for non-stationary markov decision processes: The blessing of (more) optimism
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2020
Cited alongside, same era.
Minimax-optimal off-policy evaluation with linear function approximation
Yaqi Duan, Zeyu Jia, and Mengdi Wang · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Michael Janner, Qiyang Li, and Sergey Levine · 2021
Closest in time.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Closest in time.
Nearly horizon-free offline reinforcement learning
Tongzheng Ren, Jialian Li, Bo Dai, Simon S Du, and Sujay Sanghavi · 2021
Closest in time.
Pessimistic model-based offline rl: Pac bounds and posterior sampling under partial coverage
Masatoshi Uehara and Wen Sun · 2021
Closest in time.
What are the statistical limits of offline rl with linear function approximation?
Ruosong Wang, Dean P Foster, and Sham M Kakade · 2021
Closest in time.
On the optimality of batch policy optimization algorithms
Chenjun Xiao, Yifan Wu, Jincheng Mei, Bo Dai, Tor Lattimore, Lihong Li, Csaba Szepesvari, and Dale Schuurmans · 2021
Closest in time.
Batch value-function approximation with only realizability
Tengyang Xie and Nan Jiang · 2021
Closest in time.
Optimal uniform ope and model-based offline reinforcement learning in time-homogeneous, reward-free and task-agnostic settings
Ming Yin and Yu-Xiang Wang · 2021
Closest in time.
Exponential lower bounds for batch reinforcement learning: Batch rl can be exponentially harder than online rl
Andrea Zanette · 2021
Closest in time.
Provable benefits of actor-critic methods for offline reinforcement learning
Andrea Zanette, Martin J Wainwright, and Emma Brunskill · 2021
Closest in time.
Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon
Zihan Zhang, Xiangyang Ji, and Simon S Du · 2021
Closest in time.