Fetching the paper…
Reading the bibliography…
We consider the batch (off-line) policy learning problem in the infinite horizon Markov Decision Process.
1904
Earlier work this paper cites.
1908
Earlier work this paper cites.
1909
Earlier work this paper cites.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
1911
Earlier work this paper cites.
1912
Earlier work this paper cites.
[author] Liao, P.P., Klasjna, P.P., Tewari, A.A. and Murphy, S. A.S. A. (2016). Micro-Randomized Trials in mHealth. Statistics in Medicine 35 1944-71
1944
Earlier work this paper cites.
[author] Liu, Dong CD. C. and Nocedal, JorgeJ. (1989). On the limited memory BFGS method for large scale optimization. Mathematical programming 45 503–528
1989
Earlier work this paper cites.
[author] Newey, Whitney KW. K. (1990). Semiparametric efficiency bounds. Journal of applied econometrics 5 99–135
1990
Earlier work this paper cites.
[author] Bickel, Peter JP. J., Klaassen, Chris AJC. A., Bickel, Peter JP. J., Ritov, Ya’acovY., Klaassen, JJ., Wellner, Jon AJ. A. and Ritov, YA’AcovY. (1993). Efficient and adaptive estimation for semiparametric models 4. Johns Hopkins University Press Baltimore
1993
Earlier work this paper cites.
[author] Puterman, Martin LM. L. (1994). Markov Decision Processes: Discrete Stochastic Dynamic Programming
1994
Earlier work this paper cites.
[author] Robins, James MJ. M., Rotnitzky, AndreaA. and Zhao, Lue PingL. P. (1994). Estimation of regression coefficients when some regressors are not always observed. Journal of the American statistical Association 89 846–866
1994
Earlier work this paper cites.
[author] Richardson, George BG. B. (1995). The theory of the market economy. Revue économique 1487–1496
1995
Earlier work this paper cites.
[author] Mahadevan, SridharS. (1996). Average reward reinforcement learning: Foundations, algorithms, and empirical results. Machine learning 22 159–195
1996
Earlier work this paper cites.
[author] Van Roy, BenjaminB. (1998). Learning and value function approximation in complex decision processes, PhD thesis, Massachusetts Institute of Technology
1998
Earlier work this paper cites.
[author] Hernández-Lerma, OnésimoO. and Lasserre, Jean BJ. B. (1999). Further topics on discrete-time Markov control processes 42. Springer
1999
Earlier work this paper cites.
[author] Precup, DoinaD. (2000). Eligibility traces for off-policy policy evaluation. Computer Science Department Faculty Publication Series 80
2000
Earlier work this paper cites.
[author] Van der Vaart, Aad WA. W. (2000). Asymptotic statistics 3. Cambridge university press
2000
Earlier work this paper cites.
[author] Abounadi, JinaneJ., Bertsekas, DimitribD. and Borkar, Vivek SV. S. (2001). Learning algorithms for Markov decision processes with average cost. SIAM Journal on Control and Optimization 40 681–698
2001
Earlier work this paper cites.
[author] Friedman, JeromeJ., Hastie, TrevorT. and Tibshirani, RobertR. (2001). The elements of statistical learning 1. Springer series in statistics New York
2001
Earlier work this paper cites.
[author] Murphy, Susan AS. A., van der Laan, Mark JM. J., Robins, James MJ. M. and Group, Conduct Problems Prevention ResearchC. P. P. R. (2001). Marginal mean models for dynamic regimes. Journal of the American Statistical Association 96 1410–1423
2001
Earlier work this paper cites.
2001
Earlier work this paper cites.
Kakade, S
2002
Cited alongside, same era.
[author] Lagoudakis, Michail GM. G. and Parr, RonaldR. (2003). Least-squares policy iteration. Journal of machine learning research 4 1107–1149
2003
Cited alongside, same era.
Ormoneit, D
2003
Cited alongside, same era.
[author] Ernst, DamienD., Geurts, PierreP., Wehenkel, LouisL. and Littman, L.L. (2005). Tree-based batch mode reinforcement learning. Journal of Machine Learning Research 6 503–556
2005
Cited alongside, same era.
[author] Mitrophanov, A YuA. Y. (2005). Sensitivity and convergence of uniformly ergodic Markov chains. Journal of Applied Probability 42 1003–1014
2005
Cited alongside, same era.
[author] Györfi, LászlóL., Kohler, MichaelM., Krzyzak, AdamA. and Walk, HarroH. (2006). A distribution-free theory of nonparametric regression. Springer Science & Business Media
2016
Later among the works it cites.
[author] Nahum-Shani, InbalI., Smith, Shawna NS. N., Spring, Bonnie JB. J., Collins, Linda ML. M., Witkiewitz, KatieK., Tewari, AmbujA. and Murphy, Susan AS. A. (2016). Just-in-Time Adaptive Interventions (JITAIs) in mobile health: key components and design principles for ongoing health behavior support. Annals of Behavioral Medicine 1–17
2016
Later among the works it cites.
Thomas, P
2016
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2006
Cited alongside, same era.
[author] Kosorok, Michael RM. R. (2007). Introduction to empirical processes and semiparametric inference. Springer Science & Business Media
2007
Cited alongside, same era.
[author] Tsiatis, AnastasiosA. (2007). Semiparametric theory and missing data. Springer Science & Business Media
2007
Cited alongside, same era.
[author] Munos, RémiR. and Szepesvári, CsabaC. (2008). Finite-time bounds for fitted value iteration. Journal of Machine Learning Research 9 815–857
2008
Cited alongside, same era.
[author] Steinwart, IngoI. and Christmann, AndreasA. (2008). Support vector machines. Springer Science & Business Media
2008
Cited alongside, same era.
Fukumizu, K
2009
Cited alongside, same era.
[author] Farahmand, Amir-massoudA.-m. and Szepesvári, CsabaC. (2011). Model selection in reinforcement learning. Machine learning 85 299–332
2011
Cited alongside, same era.
[author] Loh, Po-LingP.-L. et al. (2017). Statistical consistency and asymptotic normality for high-dimensional robust M M -estimators. The Annals of Statistics 45 866–896
2017
Later among the works it cites.
[author] Zhou, XinX., Mayer-Hamblett, NicoleN., Khan, UmerU. and Kosorok, Michael RM. R. (2017). Residual weighted learning for estimating individualized treatment rules. Journal of the American Statistical Association 112 169–187
2017
Later among the works it cites.
[author] Chernozhukov, VictorV., Chetverikov, DenisD., Demirer, MertM., Duflo, EstherE., Hansen, ChristianC., Newey, WhitneyW. and Robins, JamesJ. (2018). Double/debiased machine learning for treatment and structural parameters
2018
Later among the works it cites.
[author] Ertefaie, AshkanA. and Strawderman, Robert LR. L. (2018). Constructing dynamic treatment regimes over indefinite time horizons. Biometrika 105 963–977
2018
Later among the works it cites.
[author] Klasnja, PredragP., Smith, ShawnaS., Seewald, Nicholas JN. J., Lee, AndyA., Hall, KellyK., Luers, BrookB., Hekler, Eric BE. B. and Murphy, Susan AS. A. (2018). Efficacy of contextually tailored suggestions for physical activity: a micro-randomized optimization trial of HeartSteps. Annals of Behavioral Medicine
2018
Later among the works it cites.
[author] Mei, SongS., Bai, YuY., Montanari, AndreaA. et al. (2018). The landscape of empirical risk for nonconvex losses. The Annals of Statistics 46 2747–2774
2018
Later among the works it cites.
[author] Sutton, Richard SR. S. and Barto, Andrew GA. G. (2018). Reinforcement learning: An introduction. MIT press
2018
Later among the works it cites.
[author] Kosorok, Michael RM. R. and Laber, Eric BE. B. (2019). Precision medicine. Annual review of statistics and its application 6 263–286
2019
Later among the works it cites.
[author] Kumar, AviralA., Fu, JustinJ., Soh, MatthewM., Tucker, GeorgeG. and Levine, SergeyS. (2019). Stabilizing off-policy q-learning via bootstrapping error reduction. Advances in Neural Information Processing Systems 32
2019
Later among the works it cites.
[author] Luckett, Daniel JD. J., Laber, Eric BE. B., Kahkoska, Anna RA. R., Maahs, David MD. M., Mayer-Davis, ElizabethE. and Kosorok, Michael RM. R. (2019). Estimating dynamic treatment regimes in mobile health using V-learning. Journal of the American Statistical Association just-accepted 1–39
2019
Later among the works it cites.
Nachum, O
2019
Later among the works it cites.
[author] Zhao, Ying-QiY.-Q., Laber, Eric BE. B., Ning, YangY., Saha, SumonaS. and Sands, Bruce EB. E. (2019). Efficient augmentation and relaxation learning for individualized treatment rules using observational data. Journal of Machine Learning Research 20 1–23
2019
Later among the works it cites.
Agarwal, R
2020
Closest in time.
[author] Kallus, NathanN. and Uehara, MasatoshiM. (2020). Double reinforcement learning for efficient off-policy evaluation in markov decision processes. Journal of Machine Learning Research 21 1–63
2020
Closest in time.
Sharma, H
2020
Closest in time.
[author] Wu, YunanY. and Wang, LanL. (2020). Resampling-based confidence intervals for model-free robust inference on optimal treatment regimes. Biometrics n/a. https://doi.org/10.1111/biom.13337
2020
Closest in time.
Zhang, R
2020
Closest in time.
2021
Closest in time.
Fujimoto, S
2062
Closest in time.