Fetching the paper…
Reading the bibliography…
This paper presents a systematic study on gap-dependent sample complexity in offline reinforcement learning.
Marginal mean models for dynamic regimes
Susan A Murphy, Mark J van der Laan, James M Robins, and Conduct Problems Prevention Research Group · 2001
Earlier work this paper cites.
Finite time bounds for sampling based fitted value iteration
Csaba Szepesvári and Rémi Munos · 2005
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Offline policy evaluation across representations with applications to educational games
Travis Mandel, Yun-En Liu, Sergey Levine, Emma Brunskill, and Zoran Popovic · 2014
Earlier work this paper cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Top-k off-policy correction for a reinforce recommender system
Minmin Chen, Alex Beutel, Paul Covington, Sagar Jain, Francois Belletti, and Ed H Chi · 2019
Earlier work this paper cites.
Approximating interactive human evaluation with self-play for open-domain dialog systems
Asma Ghandeharioun, Judy Hanwen Shen, Natasha Jaques, Craig Ferguson, Noah Jones, Agata Lapedriza, and Rosalind Picard · 2019
Earlier work this paper cites.
Guidelines for reinforcement learning in healthcare
Omer Gottesman, Fredrik Johansson, Matthieu Komorowski, Aldo Faisal, David Sontag, Finale Doshi-Velez, and Leo Anthony Celi · 2019
Earlier work this paper cites.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Max Simchowitz and Kevin G Jamieson · 2019
Earlier work this paper cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang · 2019
Earlier work this paper cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Earlier work this paper cites.
Almost horizon-free structure-aware best policy identification with a generative model
Andrea Zanette, Mykel J Kochenderfer, and Emma Brunskill · 2019
Cited alongside, same era.
Planning in markov decision processes with gap-dependent sample complexity
Anders Jonsson, Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues, Edouard Leurent, and Michal Valko · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Cited alongside, same era.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Ming Yin and Yu-Xiang Wang · 2020
Cited alongside, same era.
Near-optimal provable uniform convergence in offline policy evaluation for reinforcement learning
Ming Yin, Yu Bai, and Yu-Xiang Wang · 2020
Cited alongside, same era.
A fully problem-dependent regret lower bound for finite-horizon mdps
Andrea Tirinzoni, Matteo Pirotta, and Alessandro Lazaric · 2021
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Masatoshi Uehara and Wen Sun · 2021
Later among the works it cites.
Representation learning for online and offline rl in low-rank mdps
Masatoshi Uehara, Xuezhou Zhang, and Wen Sun · 2021
Later among the works it cites.
Batch value-function approximation with only realizability
Tengyang Xie and Nan Jiang · 2021
Later among the works it cites.
Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Tengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong, and Yu Bai · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning
Christoph Dann, Teodor Vanislavov Marinov, Mehryar Mohri, and Julian Zimmert · 2021
Cited alongside, same era.
Logarithmic regret for reinforcement learning with linear function approximation
Jiafan He, Dongruo Zhou, and Quanquan Gu · 2021
Cited alongside, same era.
Fast rates for the regret of offline reinforcement learning, 2021
Yichun Hu, Nathan Kallus, and Masatoshi Uehara · 2021
Cited alongside, same era.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
Cited alongside, same era.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Cited alongside, same era.
Nearly horizon-free offline reinforcement learning
Tongzheng Ren, Jialian Li, Bo Dai, Simon S Du, and Sujay Sanghavi · 2021
Cited alongside, same era.
Beyond no regret: Instance-dependent pac reinforcement learning
Andrew Wagenmaker, Max Simchowitz, and Kevin Jamieson
Cited in the paper.
Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap
Haike Xu, Tengyu Ma, and Simon Du · 2021
Later among the works it cites.
Towards instance-optimal offline reinforcement learning with pessimism
Ming Yin and Yu-Xiang Wang · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Andrea Zanette, Martin J Wainwright, and Emma Brunskill · 2021
Later among the works it cites.
Settling the sample complexity of model-based offline reinforcement learning
Gen Li, Laixi Shi, Yuxin Chen, Yuejie Chi, and Yuting Wei · 2022
Closest in time.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
Laixi Shi, Gen Li, Yuting Wei, Yuxin Chen, and Yuejie Chi · 2022
Closest in time.
The efficacy of pessimism in asynchronous q-learning
Yuling Yan, Gen Li, Yuxin Chen, and Jianqing Fan · 2022
Closest in time.