Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) leverages previously collected data for policy optimization without any further active exploration.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 1905
Earlier work this paper cites.
Theory of function spaces
H. Triebel · 1983
Earlier work this paper cites.
Neuro-dynamic programming: an overview
Dimitri P Bertsekas and John N Tsitsiklis · 1995
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S. Sutton, and Satinder P. Singh · 2000
Earlier work this paper cites.
A Distribution-Free Theory of Nonparametric Regression
László Györfi, Michael Kohler, Adam Krzyzak, and Harro Walk · 2002
Earlier work this paper cites.
Least-squares policy iteration
Michail G. Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Rémi Munos · 2003
Earlier work this paper cites.
Local rademacher complexities
Peter L. Bartlett, Olivier Bousquet, and Shahar Mendelson · 2005
Earlier work this paper cites.
Model-based function approximation in reinforcement learning
Nicholas K. Jong and Peter Stone · 2007
Earlier work this paper cites.
Bracketing metric entropy rates and empirical central limit theorems for function classes of besov- and sobolev-type
R. Nickl and B. M. Pötscher · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
A primer on besov spaces, 2009
Albert Cohen · 2009
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
Modelling transition dynamics in mdps with RKHS embeddings
Steffen Grünewälder, Guy Lever, Luca Baldassarre, Massimiliano Pontil, and Arthur Gretton · 2012
Earlier work this paper cites.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning, 2015
Nan Jiang and Lihong Li · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Mathematical foundations of infinite-dimensional statistical models , volume 40
Evarist Giné and Richard Nickl · 2016
Earlier work this paper cites.
Local rademacher complexity bounds based on covering numbers
Yunwen Lei, Lixin Ding, and Yingzhou Bi · 2016
Earlier work this paper cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Earlier work this paper cites.
Learning sparse neural networks through l 0 {}_{\mbox{0}} regularization
Christos Louizos, Max Welling, and Diederik P. Kingma · 2017
Cited alongside, same era.
Boosted fitted q-iteration
Samuele Tosatto, Matteo Pirotta, Carlo D’Eramo, and Marcello Restelli · 2017
Cited alongside, same era.
Error bounds for approximations with deep relu networks
Dmitry Yarotsky · 2017
Cited alongside, same era.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Cited alongside, same era.
Max H Farrell, Tengyuan Liang, and Sanjog Misra · 2018
Cited alongside, same era.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Ming Yin and Yu-Xiang Wang · 2020
Later among the works it cites.
Mikhail Belkin · 2021
Closest in time.
Lin Chen, Bruno Scherrer, and Peter L Bartlett · 2021
Closest in time.
Fast rates for the regret of offline reinforcement learning
Yichun Hu, Nathan Kallus, and Masatoshi Uehara · 2021
Closest in time.
On the proof of global convergence of gradient descent for deep relu networks with linear widths
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gaussian processes and kernel methods: A review on connections and equivalences
Motonobu Kanagawa, Philipp Hennig, Dino Sejdinovic, and Bharath K Sriperumbudur · 2018
Cited alongside, same era.
Theory of Besov Spaces , volume 56
Yoshihiro Sawano · 2018
Cited alongside, same era.
Taiji Suzuki · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Cited alongside, same era.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 2019
Cited alongside, same era.
Quynh Nguyen · 2021
Closest in time.
Sample complexity of offline reinforcement learning with deep relu networks, 2021
Thanh Nguyen-Tang, Sunil Gupta, Hung Tran-The, and Svetha Venkatesh · 2021
Closest in time.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Closest in time.
Nearly horizon-free offline reinforcement learning
Tongzheng Ren, Jialian Li, Bo Dai, Simon S Du, and Sujay Sanghavi · 2021
Closest in time.
Pessimistic model-based offline reinforcement learning under partial coverage
Masatoshi Uehara and Wen Sun · 2021
Closest in time.
Masatoshi Uehara, Masaaki Imaizumi, Nan Jiang, Nathan Kallus, Wen Sun, and Tengyang Xie · 2021
Closest in time.
Near-optimal provable uniform convergence in offline policy evaluation for reinforcement learning
Ming Yin, Yu Bai, and Yu-Xiang Wang · 2021
Closest in time.
Offline reinforcement learning under value and density-ratio realizability: the power of gaps
Jinglin Chen and Nan Jiang · 2022
Closest in time.
Xiang Ji, Minshuo Chen, Mengdi Wang, and Tuo Zhao · 2022
Closest in time.
Understanding deep neural function approximation in reinforcement learning via $\epsilon$-greedy exploration
Fanghui Liu, Luca Viano, and Volkan Cevher · 2022
Closest in time.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
Laixi Shi, Gen Li, Yuting Wei, Yuxin Chen, and Yuejie Chi · 2022
Closest in time.
Trajectory-dependent generalization bounds for deep neural networks via fractional brownian motion
Chengli Tan, Jiangshe Zhang, and Junmin Liu · 2022
Closest in time.
Learning fractional white noises in neural stochastic differential equations
Anh Tong, Thanh Nguyen-Tang, Toan Tran, and Jaesik Choi · 2022
Closest in time.
On gap-dependent bounds for offline reinforcement learning
Xinqi Wang, Qiwen Cui, and Simon S Du · 2022
Closest in time.
Wei Xiong, Han Zhong, Chengshuai Shi, Cong Shen, Liwei Wang, and T. Zhang · 2022
Closest in time.
Ming Yin, Yaqi Duan, Mengdi Wang, and Yu-Xiang Wang · 2022
Closest in time.
Offline reinforcement learning with realizability and single-policy concentrability
Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang, and Jason D Lee · 2022
Closest in time.
Two-stage neural contextual bandits for personalised news recommendation
Mengyan Zhang, Thanh Nguyen-Tang, Fangzhao Wu, Zhenyu He, Xing Xie, and Cheng Soon Ong · 2022
Closest in time.