Fetching the paper…
Reading the bibliography…
Sample-efficient offline reinforcement learning (RL) with linear function approximation has recently been studied extensively.
On the generalization ability of on-line learning algorithms
Nicolò Cesa-Bianchi, Alex Conconi, and Claudio Gentile · 2004
Earlier work this paper cites.
Local rademacher complexities
Peter L Bartlett, Olivier Bousquet, and Shahar Mendelson · 2005
Earlier work this paper cites.
Finite time bounds for sampling based fitted value iteration
Csaba Szepesvári and Rémi Munos · 2005
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Learning from logged implicit exploration data
Alex Strehl, John Langford, Sham Kakade, and Lihong Li · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Freedman’s inequality for matrix martingales
Joel Tropp · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing
Philip S. Thomas, Georgios Theocharous, Mohammad Ghavamzadeh, Ishan Durugkar, and Emma Brunskill · 2017
Earlier work this paper cites.
Who should be treated? empirical welfare maximization methods for treatment choice
Toru Kitagawa and Aleksey Tetenov · 2018
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Earlier work this paper cites.
Guidelines for reinforcement learning in healthcare
Omer Gottesman, Fredrik Johansson, Matthieu Komorowski, Aldo Faisal, David Sontag, Finale Doshi-Velez, and Leo Anthony Celi · 2019
Earlier work this paper cites.
Batch policy learning under constraints
Hoang Minh Le, Cameron Voloshin, and Yisong Yue · 2019
Earlier work this paper cites.
Off-policy policy gradient with stationary distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Earlier work this paper cites.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Max Simchowitz and Kevin G Jamieson · 2019
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Cited alongside, same era.
Minimax-optimal off-policy evaluation with linear function approximation
Yaqi Duan, Zeyu Jia, and Mengdi Wang · 2020
Cited alongside, same era.
Learning when-to-treat policies
Xinkun Nie, Emma Brunskill, and Stefan Wager · 2021
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Masatoshi Uehara and Wen Sun · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro, and Alekh Agarwal · 2021
Later among the works it cites.
Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap
Haike Xu, Tengyu Ma, and Simon Du · 2021
Later among the works it cites.
Q-learning with logarithmic regret
Kunhe Yang, Lin Yang, and Simon Du · 2021
Later among the works it cites.
Towards instance-optimal offline reinforcement learning with pessimism
Ming Yin and Yu-Xiang Wang · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Cited alongside, same era.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Cited alongside, same era.
What are the statistical limits of offline rl with linear function approximation?
Ruosong Wang, Dean P Foster, and Sham M Kakade · 2020
Cited alongside, same era.
Policy learning with observational data
Susan Athey and Stefan Wager · 2021
Cited alongside, same era.
Logarithmic regret for reinforcement learning with linear function approximation
Jiafan He, Dongruo Zhou, and Quanquan Gu · 2021
Cited alongside, same era.
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Andrea Zanette, Martin J Wainwright, and Emma Brunskill · 2021
Later among the works it cites.
Statistical inference with m-estimators on adaptively collected data
Kelly Zhang, Lucas Janson, and Susan Murphy · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Dongruo Zhou, Quanquan Gu, and Csaba Szepesvari · 2021
Later among the works it cites.
Adversarially trained actor critic for offline reinforcement learning
Ching-An Cheng, Tengyang Xie, Nan Jiang, and Alekh Agarwal · 2022
Closest in time.
Xiang Ji, Minshuo Chen, Mengdi Wang, and Tuo Zhao · 2022
Closest in time.
On gap-dependent bounds for offline reinforcement learning
Xinqi Wang, Qiwen Cui, and Simon S Du · 2022
Closest in time.
Wei Xiong, Han Zhong, Chengshuai Shi, Cong Shen, Liwei Wang, and T. Zhang · 2022
Closest in time.
Near-optimal offline reinforcement learning with linear representation: Leveraging variance information with pessimism
Ming Yin, Yaqi Duan, Mengdi Wang, and Yu-Xiang Wang · 2022
Closest in time.