Fetching the paper…
Reading the bibliography…
Bootstrapping provides a flexible and effective approach for assessing the quality of batch reinforcement learning, yet its theoretical property is less understood.
Dependent central limit theorems and invariance principles
McLeish, D. L. et al · 1974
Earlier work this paper cites.
Some asymptotic theory for the bootstrap
Bickel, P. J. and Freedman, D. A · 1981
Earlier work this paper cites.
Bootstrapping regression models
Freedman, D. A. et al · 1981
Earlier work this paper cites.
On the asymptotic accuracy of efron’s bootstrap
Singh, K · 1981
Earlier work this paper cites.
The jackknife, the bootstrap and other resampling plans
Efron, B · 1982
Earlier work this paper cites.
Efficient memory-based learning for robot control
Moore, A. W · 1990
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R. S., and Singh, S · 2000
Earlier work this paper cites.
Asymptotic statistics , volume 3
Van der Vaart, A. W · 2000
Earlier work this paper cites.
GradientDICE: Rethinking generalized offline estimation of stationary values
Zhang, S., Liu, B., and Whiteson, S · 2001
Earlier work this paper cites.
GenDICE: Generalized offline estimation of stationary values
Zhang, R., Dai, B., Li, L., and Schuurmans, D · 2002
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Earlier work this paper cites.
The matrix cookbook. technical university of denmark
Petersen, K. and Pedersen, M · 2008
Earlier work this paper cites.
Sparse feature selection makes batch reinforcement learning more sample efficient
Hao, B., Duan, Y., Lattimore, T., Szepesvári, C., and Wang, M · 2011
Earlier work this paper cites.
A note on moment convergence of bootstrap m-estimators
Kato, K · 2011
Earlier work this paper cites.
Batch mode reinforcement learning based on the synthesis of artificial trajectories
Fonteneau, R., Murphy, S. A., Wehenkel, L., and Ernst, D · 2013
Cited alongside, same era.
A scalable bootstrap for massive data
Kleiner, A., Talwalkar, A., Sarkar, P., and Jordan, M. I · 2014
Cited alongside, same era.
High confidence policy improvement
Thomas, P., Theocharous, G., and Ghavamzadeh, M · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2016
Cited alongside, same era.
A subsampled double bootstrap for massive data
Sengupta, S., Volgushev, S., and Shao, X · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E · 2016
Cited alongside, same era.
Minimax weight and Q-function learning for off-policy evaluation
Uehara, M. and Jiang, N · 2019
Later among the works it cites.
Empirical study of off-policy policy evaluation for reinforcement learning
Voloshin, C., Le, H. M., Jiang, N., and Yue, Y · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T., Ma, Y., and Wang, Y.-X · 2019
Later among the works it cites.
Coindice: Off-policy confidence interval estimation
Dai, B., Nachum, O., Chow, Y., Li, L., Szepesvári, C., and Schuurmans, D · 2020
Later among the works it cites.
Minimax-optimal off-policy evaluation with linear function approximation
Duan, Y. and Wang, M · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Consistent on-line off-policy evaluation
Hallak, A. and Mannor, S · 2017
Cited alongside, same era.
Bootstrapping with models: Confidence intervals for off-policy evaluation
Hanna, J. P., Stone, P., and Niekum, S · 2017
Cited alongside, same era.
Bootstrapping for multivariate linear regression models
Eck, D. J · 2018
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., and Zhou, D · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Batch policy learning under constraints
Le, H. M., Voloshin, C., and Yue, Y · 2019
Cited alongside, same era.
Later among the works it cites.
Accountable off-policy evaluation with kernel bellman statistics
Feng, Y., Ren, T., Tang, Z., and Liu, Q · 2020
Later among the works it cites.
Minimax value interval for off-policy evaluation and policy optimization
Jiang, N. and Huang, J · 2020
Later among the works it cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Kallus, N. and Uehara, M · 2020
Later among the works it cites.
Statistical bootstrapping for uncertainty estimation in off-policy evaluation
Kostrikov, I. and Nachum, O · 2020
Later among the works it cites.
Confident off-policy evaluation and selection through self-normalized importance weighting
Kuzborskij, I., Vernade, C., György, A., and Szepesvári, C · 2020
Later among the works it cites.
Hyperparameter selection for offline reinforcement learning
Paine, T. L., Paduraru, C., Michi, A., Gulcehre, C., Zolna, K., Novikov, A., Wang, Z., and de Freitas, N · 2020
Later among the works it cites.
Statistical inference of the value function for reinforcement learning in infinite horizon settings
Shi, C., Zhang, S., Lu, W., and Song, R · 2020
Later among the works it cites.
What are the statistical limits of offline rl with linear function approximation?
Wang, R., Foster, D. P., and Kakade, S. M · 2020
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Yin, M. and Wang, Y.-X · 2020
Later among the works it cites.
Non-asymptotic confidence intervals of off-policy evaluation: Primal and dual bounds
Feng, Y., Tang, Z., Zhang, N., and Liu, Q · 2021
Closest in time.