Fetching the paper…
Reading the bibliography…
In reinforcement learning, it is typical to use the empirically observed transitions and rewards to estimate the value of a policy via either model-based or Q-fitting approaches.
Better bootstrap confidence intervals
Bradley Efron · 1987
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L Puterman · 1994
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Steven J Bradtke and Andrew G Barto · 1996
Earlier work this paper cites.
Bootstrap confidence intervals
Thomas J DiCiccio and Bradley Efron · 1996
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Susan A Murphy, Mark J van der Laan, James M Robins, and Conduct Problems Prevention Research Group · 2001
Earlier work this paper cites.
All of nonparametric statistics
Larry Wasserman · 2006
Earlier work this paper cites.
On the failure of the bootstrap for matching estimators
Alberto Abadie and Guido W Imbens · 2008
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Earlier work this paper cites.
Subsampling
Dimitris N Politis, Joseph P Romano, and Michael Wolf · 2012
Earlier work this paper cites.
Resampling: consistency of substitution estimators
Hein Putter and Willem R Van Zwet · 2012
Earlier work this paper cites.
The bootstrap and Edgeworth expansion
Peter Hall · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Offline policy evaluation across representations with applications to educational games
Travis Mandel, Yun-En Liu, Sergey Levine, Emma Brunskill, and Zoran Popovic · 2014
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
High confidence off-policy evaluation
P. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Cited alongside, same era.
Off-policy evaluation for slate recommendation
A. Swaminathan, A. Krishnamurthy, A. Agarwal, M. Dudík, J. Langford, D. Jose, and I. Zitouni · 2017
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Garbage in, reward out: Bootstrapping exploration in multi-armed bandits
Branislav Kveton, Csaba Szepesvari, Sharan Vaswani, Zheng Wen, Mohammad Ghavamzadeh, and Tor Lattimore · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Later among the works it cites.
Bootstrapping upper confidence bound
Botao Hao, Yasin Abbasi Yadkori, Zheng Wen, and Guang Cheng · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
High confidence policy improvement
P. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Cited alongside, same era.
Safe reinforcement learning
Philip S Thomas · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
High confidence off-policy evaluation with models
J. Hanna, P. Stone, and S. Niekum · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
P. Thomas and E. Brunskill · 2016
Cited alongside, same era.
Bootstrapping with models: Confidence intervals for off-policy evaluation
Josiah P Hanna, Peter Stone, and Scott Niekum · 2017
Cited alongside, same era.
Perturbed-history exploration in stochastic multi-armed bandits
Branislav Kveton, Csaba Szepesvari, Mohammad Ghavamzadeh, and Craig Boutilier · 2019
Later among the works it cites.
Off-policy estimation of long-term average outcomes with applications to mobile health
Peng Liao, Predrag Klasnja, and Susan Murphy · 2019
Later among the works it cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Later among the works it cites.
Why does hierarchy (sometimes) work so well in reinforcement learning?
Ofir Nachum, Haoran Tang, Xingyu Lu, Shixiang Gu, Honglak Lee, and Sergey Levine · 2019
Later among the works it cites.
Empirical study of off-policy policy evaluation for reinforcement learning
Cameron Voloshin, Hoang M Le, Nan Jiang, and Yisong Yue · 2019
Later among the works it cites.
Minimax-optimal off-policy evaluation with linear function approximation
Yaqi Duan and Mengdi Wang · 2020
Closest in time.
Reinforcement learning via fenchel-rockafellar duality
Ofir Nachum and Bo Dai · 2020
Closest in time.
Hyperparameter selection for offline reinforcement learning, 2020
Tom Le Paine, Cosmin Paduraru, Andrea Michi, Caglar Gulcehre, Konrad Zolna, Alexander Novikov, Ziyu Wang, and Nando de Freitas · 2020
Closest in time.