Fetching the paper…
Reading the bibliography…
Reinforcement learning is a general technique that allows an agent to learn an optimal policy and interact with an environment in sequential decision making problems.
arXiv preprint arXiv:1904.07103
Meitz, M. and Saikkonen, P. (2019) Subgeometric ergodicity and β \beta -mixing · 1904
Earlier work this paper cites.
arXiv preprint arXiv:1909.05850
Kallus, N. and Uehara, M. (2019) Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning · 1909
Earlier work this paper cites.
Teoriya Veroyatnostei i ee Primeneniya
Davydov, Y. A. (1973) Mixing conditions for markov chains · 1973
Earlier work this paper cites.
Ann. Probability
McLeish, D. L. (1974) Dependent central limit theorems and invariance principles · 1974
Earlier work this paper cites.
John Wiley & Sons, Inc., New York
Schumaker, L. L. (1981) Spline functions: basic theory · 1981
Earlier work this paper cites.
Probab. Surv
Bradley, R. C. (2005) Basic properties of strong mixing conditions. A survey and some open questions · 1986
Earlier work this paper cites.
Ann. Statist
Burman, P. and Chen, K.-W. (1989) Nonparametric estimation of a regression function · 1989
Earlier work this paper cites.
Cambridge University Press, Cambridge
Meyer, Y. (1992) Wavelets and operators · 1990
Earlier work this paper cites.
Wiley Series in Probability and Mathematical Statistics: Applied Probability and Statistics. John Wiley & Sons, Inc., New York
Puterman, M. L. (1994) Markov decision processes: discrete stochastic dynamic programming · 1994
Earlier work this paper cites.
Ann. Statist
Huang, J. Z. (1998) Projection estimation in multiple regression with application to functional ANOVA models · 1998
Earlier work this paper cites.
Tech. rep
Saikkonen, P. (2001) Stability results for nonlinear vector autoregressions with an application to a nonlinear error correction model · 2001
Earlier work this paper cites.
J. R. Stat. Soc. Ser. B Stat. Methodol
Murphy, S. A. (2003) Optimal dynamic treatment regimes · 2003
Earlier work this paper cites.
Ann. Statist
Tsybakov, A. B. (2004) Optimal aggregation of classifiers in statistical learning · 2004
Earlier work this paper cites.
J. Mach. Learn. Res
Ernst, D., Geurts, P. and Wehenkel, L. (2005) Tree-based batch mode reinforcement learning · 2005
Earlier work this paper cites.
In European Conference on Machine Learning
Riedmiller, M. (2005) Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method · 2005
Earlier work this paper cites.
Ann. Statist
Audibert, J.-Y. and Tsybakov, A. B. (2007) Fast learning rates for plug-in classifiers · 2007
Earlier work this paper cites.
Diabetes technology & therapeutics
Rodbard, D. (2009) Interpretation of continuous glucose monitoring data: glycemic variability and quality of glycemic control · 2009
Cited alongside, same era.
In Advances in Neural Information Processing Systems
Hasselt, H. V. (2010) Double q-learning · 2010
Cited alongside, same era.
Ann. Statist
Qian, M. and Murphy, S. A. (2011) Performance guarantees for individualized treatment rules · 2011
Cited alongside, same era.
Electron. Commun. Probab
Tropp, J. A. (2011) Freedman’s inequality for matrix martingales · 2011
Cited alongside, same era.
Found. Comput. Math
— (2012) User-friendly tail bounds for sums of random matrices · 2012
Cited alongside, same era.
Robotics
Kormushev, P., Calinon, S. and Caldwell, D. (2013) Reinforcement learning in robotics: Applications and real-world challenges · 2013
Cited alongside, same era.
Biometrika
Ertefaie, A. and Strawderman, R. L. (2018) Constructing dynamic treatment regimes over indefinite time horizons · 2018
Later among the works it cites.
In Proceedings of the 27th ACM International Conference on Information and Knowledge Management
Jin, J., Song, C., Li, H., Gai, K., Wang, J. and Zhang, W. (2018) Real-time bidding with multi-agent reinforcement learning in display advertising · 2018
Later among the works it cites.
In KHD@ IJCAI
Marling, C. and Bunescu, R. C. (2018) The ohiot1dm dataset for blood glucose level prediction · 2018
Later among the works it cites.
Adaptive Computation and Machine Learning. MIT Press, Cambridge, MA, second edn
Sutton, R. S. and Barto, A. G. (2018) Reinforcement learning: an introduction · 2018
Later among the works it cites.
In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
Xu, Z., Li, Z., Guan, Q., Zhang, D., Li, Q., Nan, J., Liu, C., Bian, W. and Ye, J. (2018) Large-scale order dispatch in on-demand ride-hailing platforms: A learning and planning approach · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, B., Tsiatis, A. A., Laber, E. B. and Davidian, M. (2013) Robust estimation of optimal dynamic treatment regimes for sequential treatment decisions · 2013
Cited alongside, same era.
J. Econometrics
Chen, X. and Christensen, T. M. (2015) Optimal uniform convergence rates and asymptotic normality for series estimators under weak dependence and weak conditions · 2015
Cited alongside, same era.
Statistical science
Dezeure, R., Bühlmann, P., Meier, L. and Meinshausen, N. (2015) High-dimensional inference: Confidence intervals, p-values and r-software hdi · 2015
Cited alongside, same era.
In Twenty-Ninth AAAI Conference on Artificial Intelligence
Thomas, P. S., Theocharous, G. and Ghavamzadeh, M. (2015) High-confidence off-policy evaluation · 2015
Cited alongside, same era.
International journal of epidemiology
Tsao, C. W. and Vasan, R. S. (2015) Cohort profile: The framingham heart study (fhs): overview of milestones in cardiovascular epidemiology · 2015
Cited alongside, same era.
J. Amer. Statist. Assoc
Zhao, Y.-Q., Zeng, D., Laber, E. B. and Kosorok, M. R. (2015) New statistical learning methods for estimating optimal dynamic treatment regimes · 2015
Cited alongside, same era.
J. Amer. Statist. Assoc
Zhang, Y., Laber, E. B., Davidian, M. and Tsiatis, A. A. (2018) Estimation of optimal treatment regimes using lists · 2018
Later among the works it cites.
In Advances in Neural Information Processing Systems
Janner, M., Fu, J., Zhang, M. and Levine, S. (2019) When to trust your model: Model-based policy optimization · 2019
Later among the works it cites.
J. Amer. Statist. Assoc
Luckett, D. J., Laber, E. B., Kahkoska, A. R., Maahs, D. M., Mayer-Davis, E. and Kosorok, M. R. (2019) Estimating dynamic treatment regimes in mobile health using v-learning · 2019
Later among the works it cites.
In Learning for Dynamics and Control
Fan, J., Wang, Z., Xie, Y. and Yang, Z. (2020) A theoretical analysis of deep q-learning · 2020
Closest in time.
Journal of Machine Learning Research
— (2020) Double reinforcement learning for efficient off-policy evaluation in markov decision processes · 2020
Closest in time.
In International Conference on Learning Representations
Tang, Z., Feng, Y., Li, L., Zhou, D. and Liu, Q. (2020) Doubly robust bias reduction in infinite horizon off-policy estimation · 2020
Closest in time.
In International Conference on Machine Learning
Uehara, M., Huang, J. and Jiang, N. (2020) Minimax weight and q-function learning for off-policy evaluation · 2020
Closest in time.
Journal of the American Statistical Association
Wang, J., He, X. and Xu, G. (2020) Debiased inference on treatment effect in a high-dimensional model · 2020
Closest in time.
arXiv preprint arXiv:2102.00479
Hu, Y., Kallus, N. and Uehara, M. (2021) Fast rates for the regret of offline reinforcement learning · 2021
Closest in time.
In International Conference on Machine Learning
Shi, C., Wan, R., Chernozhukov, V. and Song, R. (2021) Deeply-debiased off-policy interval estimation · 2021
Closest in time.