Fetching the paper…
Reading the bibliography…
This article reviews the recent advances on the statistical foundation of reinforcement learning (RL) in the offline and low-adaptive settings.
Algaedice: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 1912
Earlier work this paper cites.
Theory of statistical estimation
Ronald Aylmer Fisher · 1925
Earlier work this paper cites.
A generalization of sampling without replacement from a finite universe
Daniel G Horvitz and Donovan J Thompson · 1952
Earlier work this paper cites.
Dynamic programming
Richard Bellman · 1966
Earlier work this paper cites.
Dependent central limit theorems and invariance principles
Donald L McLeish · 1974
Earlier work this paper cites.
Maxmin expected utility with non-unique prior
Itzhak Gilboa and David Schmeidler · 1989
Earlier work this paper cites.
Markov decision processes
Martin L Puterman · 1990
Earlier work this paper cites.
Bootstrapping: A nonparametric approach to statistical inference
Christopher Z Mooney, Robert D Duval, and Robert Duvall · 1993
Earlier work this paper cites.
Approximate solutions to Markov decision processes
Geoffrey J Gordon · 1999
Earlier work this paper cites.
Degenerate nonlinear programming with a quadratic growth condition
Mihai Anitescu · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Asymptotic statistics , volume 3
Aad W Van der Vaart · 2000
Earlier work this paper cites.
Monte Carlo strategies in scientific computing , volume 10
Jun S Liu and Jun S Liu · 2001
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Susan A Murphy, Mark J van der Laan, James M Robins, and Conduct Problems Prevention Research Group · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem, 2002
P Auer · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Efficient estimation of average treatment effects using the estimated propensity score
Keisuke Hirano, Guido W Imbens, and Geert Ridder · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Finite time bounds for sampling based fitted value iteration
Csaba Szepesvári and Rémi Munos · 2005
Earlier work this paper cites.
Semiparametric theory and missing data , volume 4
Anastasios A Tsiatis · 2006
Earlier work this paper cites.
Fitted q-iteration in continuous action-space mdps
András Antos, Csaba Szepesvári, and Rémi Munos · 2007
Earlier work this paper cites.
Approximate Dynamic Programming: Solving the curses of dimensionality , volume 703
Warren B Powell · 2007
Earlier work this paper cites.
Introduction to empirical processes and semiparametric inference , volume 61
Michael R Kosorok · 2008
Earlier work this paper cites.
Minimax theory
John Lafferty, Han Liu, and Larry Wasserman · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Earlier work this paper cites.
Online learning with switching costs and other adaptive adversaries
Nicolo Cesa-Bianchi, Ofer Dekel, and Ohad Shamir · 2013
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Earlier work this paper cites.
How hard is my mdp?" the distribution-norm to the rescue"
Odalric-Ambrym Maillard, Timothy A Mann, and Shie Mannor · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Earlier work this paper cites.
Optimal medication dosing from suboptimal clinical examples: A deep reinforcement learning approach
Shamim Nemati, Mohammad M Ghassemi, and Gari D Clifford · 2016
Earlier work this paper cites.
Batched bandit problems
Vianney Perchet, Philippe Rigollet, Sylvain Chassang, and Erik Snowberg · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Consistent on-line off-policy evaluation
Assaf Hallak and Shie Mannor · 2017
Cited alongside, same era.
Neural adaptive video streaming with pensieve
Hongzi Mao, Ravi Netravali, and Mohammad Alizadeh · 2017
Cited alongside, same era.
Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods
Eyke Hüllermeier and Willem Waegeman · 2021
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
Later among the works it cites.
Offline reinforcement learning with fisher divergence critic regularization
Ilya Kostrikov, Rob Fergus, Jonathan Tompson, and Ofir Nachum · 2021
Later among the works it cites.
Deployment-efficient reinforcement learning via model-based offline optimization
Tatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum, and Shixiang Gu · 2021
Later among the works it cites.
Variance-aware off-policy evaluation with linear function approximation
Yifei Min, Tianhao Wang, Dongruo Zhou, and Quanquan Gu · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Optimal and adaptive off-policy evaluation in contextual bandits
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudık · 2017
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Deep reinforcement learning for page-wise recommendations
Xiangyu Zhao, Long Xia, Liang Zhang, Zhuoye Ding, Dawei Yin, and Jiliang Tang · 2018
Cited alongside, same era.
Provably efficient q-learning with low switching cost
Yu Bai, Tengyang Xie, Nan Jiang, and Yu-Xiang Wang · 2019
Cited alongside, same era.
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Later among the works it cites.
Nearly horizon-free offline reinforcement learning
Tongzheng Ren, Jialian Li, Bo Dai, Simon S Du, and Sujay Sanghavi · 2021
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation under adaptivity constraints
Tianhao Wang, Dongruo Zhou, and Quanquan Gu · 2021
Later among the works it cites.
On the optimality of batch policy optimization algorithms
Chenjun Xiao, Yifan Wu, Jincheng Mei, Bo Dai, Tor Lattimore, Lihong Li, Csaba Szepesvari, and Dale Schuurmans · 2021
Later among the works it cites.
Towards instance-optimal offline reinforcement learning with pessimism
Ming Yin and Yu-Xiang Wang · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Dongruo Zhou, Quanquan Gu, and Csaba Szepesvari · 2021
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Later among the works it cites.
Discovering faster matrix multiplication algorithms with reinforcement learning
Alhussein Fawzi, Matej Balog, Aja Huang, Thomas Hubert, Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Francisco J R Ruiz, Julian Schrittwieser, Grzegorz Swirszcz, et al · 2022
Later among the works it cites.
Towards deployment-efficient reinforcement learning: Lower bound and optimality
Jiawei Huang, Jinglin Chen, Li Zhao, Tao Qin, Nan Jiang, and Tie-Yan Liu · 2022
Later among the works it cites.
Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning
Nathan Kallus and Masatoshi Uehara · 2022
Later among the works it cites.
Mildly conservative q-learning for offline reinforcement learning
Jiafei Lyu, Xiaoteng Ma, Xiu Li, and Zongqing Lu · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Sample-efficient reinforcement learning with loglog (t) switching cost
Dan Qiao, Ming Yin, Ming Min, and Yu-Xiang Wang · 2022
Later among the works it cites.
On gap-dependent bounds for offline reinforcement learning
Xinqi Wang, Qiwen Cui, and Simon S Du · 2022
Later among the works it cites.
Wei Xiong, Han Zhong, Chengshuai Shi, Cong Shen, Liwei Wang, and Tong Zhang · 2022
Later among the works it cites.
Ming Yin, Yaqi Duan, Mengdi Wang, and Yu-Xiang Wang · 2022
Later among the works it cites.
Pessimistic nonlinear least-squares value iteration for offline reinforcement learning
Qiwei Di, Heyang Zhao, Jiafan He, and Quanquan Gu · 2023
Later among the works it cites.
Offline reinforcement learning with closed-form policy improvement operators
Jiachen Li, Edwin Zhang, Ming Yin, Qinxun Bai, Yu-Xiang Wang, and William Yang Wang · 2023
Later among the works it cites.
Faster sorting algorithms discovered using deep reinforcement learning
Daniel J Mankowitz, Andrea Michi, Anton Zhernov, Marco Gelmi, Marco Selvi, Cosmin Paduraru, Edouard Leurent, Shariq Iqbal, Jean-Baptiste Lespiau, Alex Ahern, et al · 2023
Later among the works it cites.
On instance-dependent bounds for offline reinforcement learning with linear function approximation
Thanh Nguyen-Tang, Ming Yin, Sunil Gupta, Svetha Venkatesh, and Raman Arora · 2023
Later among the works it cites.
Near-optimal deployment efficiency in reward-free reinforcement learning with linear function approximation
Dan Qiao and Yu-Xiang Wang · 2023
Later among the works it cites.
Offline reinforcement learning with differentiable function approximation is provably efficient
Ming Yin, Mengdi Wang, and Yu-Xiang Wang · 2023
Later among the works it cites.
Learning the target network in function space
Kavosh Asadi, Yao Liu, Shoham Sabach, Ming Yin, and Rasool Fakoor · 2024
Later among the works it cites.
Networkgym: Reinforcement learning environments for multi-access traffic management in network simulation
Momin Haider, Ming Yin, Menglei Zhang, Arpit Gupta, Jing Zhu, and Yu-Xiang Wang · 2024
Later among the works it cites.
Offline reinforcement learning in large state spaces: Algorithms and guarantees
Nan Jiang and Tengyang Xie · 2024
Later among the works it cites.
On sample-efficient offline reinforcement learning: Data diversity, posterior sampling and beyond
Thanh Nguyen-Tang and Raman Arora · 2024
Later among the works it cites.
Logarithmic switching cost in reinforcement learning beyond linear mdps
Dan Qiao, Ming Yin, and Yu-Xiang Wang · 2024
Later among the works it cites.
A nearly optimal and low-switching algorithm for reinforcement learning with general function approximation
Heyang Zhao, Jiafan He, and Quanquan Gu · 2024
Later among the works it cites.