Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) is one of the most vibrant research frontiers in machine learning and has been recently applied to solve a number of challenging problems.
1901
Earlier work this paper cites.
1908
Earlier work this paper cites.
1909
Earlier work this paper cites.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
[author] Rosenbaum, P. R.P. R. (1983). The central role of the propensity score in observational studies for causal effects. 70 41–55
1983
Earlier work this paper cites.
[author] Robins, James MJ. M., Li, LinglingL., Mukherjee, RajarshiR., Tchetgen, Eric TchetgenE. T. and van der Vaart, AadA. (2017). Minimax estimation of a functional on a structured high-dimensional model. The Annals of Statistics 45 1951–1987
1987
Earlier work this paper cites.
[author] Chamberlain, GaryG. (1992). Comment: Sequential moment restrictions in panel data. Journal of Business & Economic Statistics 10 20–26
1992
Earlier work this paper cites.
[author] Robins, James M.J. M., Rotnitzky, AndreaA. and Zhao, Lue PingL. P. (1994). Estimation of Regression Coefficients When Some Regressors are not Always Observed. Journal of the American Statistical Association 89 846–866
1994
Earlier work this paper cites.
[author] Manski, Charles FC. F. (1995). Identification problems in the social sciences. Harvard University Press, Cambridge, Mass
1995
Earlier work this paper cites.
[author] Robins, James MJ. M., Rotnitzky, AndreaA. and Zhao, Lue PingL. P. (1995). Analysis of Semiparametric Regression Models for Repeated Outcomes in the Presence of Missing Data. Journal of the American Statistical Association 90 106–121
1995
Earlier work this paper cites.
[author] Hahn, JinyongJ. (1998). On the Role of the Propensity Score in Efficient Semiparametric Estimation of Average Treatment Effects. Econometrica 66 315–331
1998
Earlier work this paper cites.
[author] Heckman, James J.J. J., Ichimura, HidehikoH. and Todd, PetraP. (1998). Matching as an econometric evaluation estimator. Review of Economic Studies 65
1998
Earlier work this paper cites.
[author] van der Vaart, A. W.A. W. (1998). Asymptotic statistics. Cambridge University Press, Cambridge, UK
1998
Earlier work this paper cites.
[author] Robins, J. M.J. M., Rotnitzky, A.A. and Scharfstein, D. O.D. O. (1999). Sensitivity Analysis for Selection Bias and Unmeasured Confounding in Missing Data and Causal Inference Models. Statistical Models in Epidemiology: The Environment and Clinical Trials. 116. NY: Springer-Verlag
1999
Earlier work this paper cites.
[author] Scharfstein, D.D., Rotnizky, A.A. and Robins, J. M.J. M. (1999). Adjusting for nonignorable dropout using semi-parametric models. Journal of the American Statistical Association 94 1096–1146
1999
Earlier work this paper cites.
[author] Precup, D.D., Sutton, R.R. and Singh, SS. (2000). Eligibility traces for off-policy policy evaluation. In Proceedings of the 17th International Conference on Machine Learning 759–766
2000
Earlier work this paper cites.
2001
Earlier work this paper cites.
[author] Murphy, S. A.S. A., van der Laan, M. J.M. J., Robins, J. M.J. M. and Group, Conduct Problems Prevention ResearchC. P. P. R. (2001). Marginal Mean Models for Dynamic Regimes. Journal of the American Statistical Association 96 1410–1423
2001
Earlier work this paper cites.
2001
Earlier work this paper cites.
[author] Rosenbaum, Paul R.P. R. (2002). Observational Studies, second edition. ed. Springer Series in Statistics. Springer New York : Imprint: Springer, New York, NY
2002
Earlier work this paper cites.
[author] Ai, ChunrongC. and Chen, XiaohongX. (2003). Efficient estimation of models with conditional moment restrictions containing unknown functions. Econometrica 71 1795–1843
2003
Earlier work this paper cites.
2003
Earlier work this paper cites.
[author] Hirano, K.K., Imbens, G.G. and Ridder, G.G. (2003). Efficient estimation of average treatment effects using the estimated propensity score. Econometrica 71 1161–1189
2003
Earlier work this paper cites.
[author] Murphy, S. A.S. A. (2003). Optimal dynamic treatment regimes. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 65 331–355
2003
Earlier work this paper cites.
2003
Earlier work this paper cites.
[author] van Der Laan, Mark J.M. J. and Robins, James MJ. M. (2003). Unified Methods for Censored Longitudinal Data and Causality. Springer Series in Statistics,. Springer New York, New York, NY
2003
Earlier work this paper cites.
[author] Lagoudakis, MichailM. and Parr, RonaldR. (2004). Least-Squares Policy Iteration. Journal of Machine Learning Research 4 1107–1149
2004
Earlier work this paper cites.
[author] Robins, J. M.J. M. (2004). Optimal structural nested models for optimal sequentialdecisions. In Proceedings of the Second Seattle Symposium in Biostatistics: Analysis of Correlated Data
2004
Earlier work this paper cites.
[author] Bang, HeejungH. and Robins, James MJ. M. (2005). Doubly robust estimation in missing data and causal inference models. Biometrics 61 962–973
2005
Earlier work this paper cites.
[author] Ernst, DamienD., Geurts, PierreP. and Wehenkel, LouisL. (2005). Tree-based batch mode reinforcement learning. Journal of Machine Learning Research 6 503–556
2005
Earlier work this paper cites.
2005
Earlier work this paper cites.
[author] Rubin, D. B.D. B. (2005). Causal inference using potential outcomes: design, modeling, decisions. Journal of the American Statistical Association 100 322-331
2005
Earlier work this paper cites.
[author] Dikkala, NishanthN., Lewis, GregG., Mackey, LesterL. and Syrgkanis, VasilisV. (2020). arXiv preprint arXiv: 2006.07201
2006
Earlier work this paper cites.
2006
Earlier work this paper cites.
2006
Earlier work this paper cites.
Shpitser, I
2006
Earlier work this paper cites.
[author] Tsiatis, Anastasios AA. A. (2006). Semiparametric Theory and Missing Data. Springer Series in Statistics. Springer New York, New York, NY
2006
Earlier work this paper cites.
Bennett, A
2007
Earlier work this paper cites.
[author] Kang, Joseph D. Y.J. D. Y. and Schafer, Joseph L.J. L. (2007). Demystifying Double Robustness: A Comparison of Alternative Strategies for Estimating a Population Mean from Incomplete Data. Statistical Science 22 523–539
2007
Earlier work this paper cites.
[author] Robins, JamesJ., Sued, MarielaM., Lei-Gomez, QuanhongQ. and Rotnitzky, AndreaA. (2007). Comment: Performance of Double-Robust Estimators When "Inverse Probability" Weights Are Highly Variable. Statistical Science 22 544–559
2007
Earlier work this paper cites.
[author] Tsiatis, Anastasios AA. A. and Davidian, MarieM. (2007). Comment: Demystifying Double Robustness: A Comparison of Alternative Strategies for Estimating a Population Mean from Incomplete Data. Statistical science 22 569–573
2007
Earlier work this paper cites.
2007
Earlier work this paper cites.
[author] Antos, AndrásA., Szepesvári, CsabaC. and Munos, RémiR. (2008). Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path. Machine Learning 71 89–129
2008
Earlier work this paper cites.
[author] Munos, RémiR. and Szepesvári, CsabaC. (2008). Finite-time bounds for fitted value iteration. Journal of Machine Learning Research 9 815–857
2008
Earlier work this paper cites.
[author] Rubin, Daniel BD. B. and van Der Laan, Mark JM. J. (2008). Empirical efficiency maximization: improved locally efficient covariate adjustment in randomized experiments and survival analysis. The international journal of biostatistics 4
2008
Earlier work this paper cites.
[author] Steinwart, IngoI. and Christmann, AndreasA. (2008). Support vector machines. Springer Science & Business Media
2008
Earlier work this paper cites.
[author] Bertsekas, Dimitri PD. P. and Yu, HuizhenH. (2009). Projected equation methods for approximate solution of large linear systems. Journal of computational and applied mathematics 227 27–50
2009
Earlier work this paper cites.
[author] Cao, WeihuaW., Tsiatis, Anastasios A.A. A. and Davidian, MarieM. (2009). Improving efficiency and robustness of the doubly robust estimator for a population mean with incomplete data. Biometrika 96 723–734
2009
Earlier work this paper cites.
[author] Meinshausen, NicolaiN., Meier, LukasL. and Bühlmann, PeterP. (2009). P-values for high-dimensional regression. Journal of the American Statistical Association 104 1671–1681
2009
Earlier work this paper cites.
[author] Tsybakov, Alexandre BA. B. (2009). Lower bounds on the minimax risk. In Introduction to Nonparametric Estimation 77–135. Springer
2009
Earlier work this paper cites.
2010
Earlier work this paper cites.
[author] Tan, ZhiqiangZ. (2010). Bounded, efficient and doubly robust estimation with inverse weighting. Biometrika 97 661–682
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
[author] Chapelle, OlivierO. and Li, LihongL. (2011). An Empirical Evaluation of Thompson Sampling. In Advances in Neural Information Processing Systems 24 2249–2257
2011
Earlier work this paper cites.
[author] Qian, MinM. and Murphy, Susan AS. A. (2011). Performance guarantees for individualized treatment rules. Annals of statistics 39 1180
2011
Earlier work this paper cites.
[author] Tsiatis, Anastasios AA. A., Davidian, MarieM. and Cao, WeihuaW. (2011). Improved Doubly Robust Estimation When Data Are Monotonely Coarsened, with Application to Longitudinal Studies with Dropout. Biometrics 67 536–545
2011
Earlier work this paper cites.
[author] Zheng, WenjingW. and van Der Laan, Mark JM. J. (2011). Cross-Validated Targeted Minimum-Loss-Based Estimation. In Targeted Learning: Causal Inference for Observational and Experimental Data. Springer Series in Statistics 459–474. Springer New York, New York, NY
2011
Earlier work this paper cites.
[author] Ai, ChunrongC. and Chen, XiaohongX. (2012). The semiparametric efficiency bound for models of sequential moment restrictions containing unknown functions. Journal of Econometrics 170 442–457
2012
Earlier work this paper cites.
[author] Bertsekas, Dimitri PD. P. (2012). Dynamic programming and optimal control, 4th ed. ed. Athena Scientific optimization and computation series. Athena Scientific, Belmont, Mass
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
[author] Wang, LuL., Rotnitzky, AndreaA., Lin, XihongX., Millikan, Randall ER. E. and Thall, Peter FP. F. (2012). Evaluation of viable dynamic treatment regimes in a sequentially randomized trial of advanced prostate cancer. Journal of the American Statistical Association 107 493–508
2012
Cited alongside, same era.
[author] Zhang, BaqunB., Tsiatis, Anastasios AA. A., Laber, Eric BE. B. and Davidian, MarieM. (2012). A robust method for estimating optimal treatment regimes. Biometrics 68 1010–1018
2012
Cited alongside, same era.
[author] Kober, JensJ., Bagnell, J AndrewJ. A. and Peters, JanJ. (2013). Reinforcement learning in robotics: A survey. The International Journal of Robotics Research 32 1238–1274
2013
Cited alongside, same era.
[author] Lu, WenbinW., Zhang, Hao HelenH. H. and Zeng, DonglinD. (2013). Variable selection for optimal treatment decision. Statistical methods in medical research 22 493–504
2013
Cited alongside, same era.
[author] Feng, YihaoY., Li, LihongL. and Liu, QiangQ. (2019). A Kernel Loss for Solving the Bellman Equation. In Advances in Neural Information Processing Systems 32 15430–15441
2019
Later among the works it cites.
[author] Gottesman, OmerO., Johansson, FredrikF., Komorowski, MatthieuM., Faisal, AldoA., Sontag, DavidD., Doshi-Velez, FinaleF. and Celi, Leo AnthonyL. A. (2019). Guidelines for reinforcement learning in healthcare. Nat Med 25 16–18
2019
Later among the works it cites.
[author] Hernan, M. A.M. A. and Robins, J. M.J. M. (2019). Causal Inference. Boca Raton: Chapman & Hall/CRC
2019
Later among the works it cites.
[author] Kennedy, Edward HE. H. (2019). Nonparametric Causal Effects Based on Incremental Propensity Score Interventions. Journal of the American Statistical Association 114 645–656
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
[author] Newey, Whitney KW. K. (2013). Nonparametric instrumental variables estimation. American Economic Review 103 550–56
2013
Cited alongside, same era.
[author] Owen, Art B.A. B. (2013). Monte Carlo theory, methods and examples
2013
Cited alongside, same era.
[author] Zhang, BaqunB., Tsiatis, Anastasios A.A. A., Laber, Eric B.E. B. and Davidian, MarieM. (2013). Robust estimation of optimal dynamic treatment regimes for sequential treatment decisions. Biometrika 100 681–694
2013
Cited alongside, same era.
[author] Chakraborty, BibhasB., Laber, Eric BE. B. and Zhao, Ying-QiY.-Q. (2014). Inference about the expected performance of a data-driven dynamic treatment regime. Clinical Trials 11 408–417
2014
Cited alongside, same era.
[author] Dann, ChristophC., Neumann, GerhardG. and Peters, JanJ. (2014). Policy Evaluation with Temporal Differences: A Survey and Comparison. Journal of Machine Learning Research 15 809-883
2014
Cited alongside, same era.
[author] Dudik, MiroslavM., Erhan, DumitruD., Langford, JohnJ. and Li, LihongL. (2014). Doubly Robust Policy Evaluation and Optimization. Statistical Science 29 485–511
2014
Cited alongside, same era.
[author] Imai, KosukeK. and Ratkovic, MarcM. (2014). Covariate balancing propensity score. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 76 243–263
2014
Cited alongside, same era.
[author] Mandel, T.T., Liu, Y.Y., Levine, S.S., Brunskill, E.E. and Popovic, ZZ. (2014). Off-policy evaluation across representations with applications to educational games. In Proceedings of the 13th International Conference on Autonomous Agentsand Multi-agent Systems 1077–1084
2014
Cited alongside, same era.
2019
Later among the works it cites.
[author] Narita, YusukeY., Yasui, ShotaS. and Yata, KoheiK. (2019). Efficient Counterfactual Learning from Bandit Feedback. AAAI
2019
Later among the works it cites.
[author] Romano, Joseph PJ. P. and DiCiccio, CyrusC. (2019). Multiple data splitting for testing. Department of Statistics, Stanford University
2019
Later among the works it cites.
[author] Tsiatis, Anastasios AA. A., Davidian, MarieM., Holloway, Shannon TS. T. and Laber, Eric BE. B. (2019). Dynamic Treatment Regimes: Statistical Methods for Precision Medicine. CRC press
2019
Later among the works it cites.
[author] Xie, TengyangT., Ma, YifeiY. and Wang, Yu-XiangY.-X. (2019). Towards Optimal Off-Policy Evaluation for Reinforcement Learning with Marginalized Importance Sampling. In Advances in Neural Information Processing Systems 32 9665–9675
2019
Later among the works it cites.
[author] Zhu, WenshengW., Zeng, DonglinD. and Song, RuiR. (2019). Proper inference for value function in high-dimensional Q-learning for dynamic treatment regimes. Journal of the American Statistical Association 114 1404–1417
2019
Later among the works it cites.
2020
Later among the works it cites.
[author] Clifton, JesseJ. and Laber, EricE. (2020). Q-Learning: Theory and Applications. Annual review of statistics and its application 7 279–301
2020
Later among the works it cites.
[author] Fulcher, Isabel RI. R., Shpitser, IlyaI., Marealle, StellaS. and Tchetgen Tchetgen, Eric JE. J. (2020). Robust inference on population indirect causal effects: the generalized front door criterion. Journal of the Royal Statistical Society. Series B, Statistical methodology 82 199–214
2020
Later among the works it cites.
[author] Hu, XinyuX., Qian, MinM., Cheng, BinB. and Cheung, Ying KuenY. K. (2020). Personalized Policy Learning using Longitudinal Mobile Health Data. Journal of the American Statistical Association 1–11
2020
Later among the works it cites.
Kallus, N
2020
Later among the works it cites.
Kidambi, R
2020
Later among the works it cites.
[author] Liao, PengP., Klasnja, PredragP. and Murphy, SusanS. (2020). Off-Policy Estimation of Long-Term Average Outcomes with Applications to Mobile Health. Journal of the American Statistical Association (To appear)
2020
Later among the works it cites.
[author] Luckett, Daniel J.D. J., Laber, Eric B.E. B., Kahkoska, Anna R.A. R., Maahs, David M.D. M., Mayer-Davis, ElizabethE. and Kosorok, Michael R.M. R. (2020). Estimating Dynamic Treatment Regimes in Mobile Health Using V-Learning. Journal of the American Statistical Association 115 692–706
2020
Later among the works it cites.
[author] Nakada, RyumeiR. and Imaizumi, MasaakiM. (2020). Adaptive Approximation and Generalization of Deep Neural Network with Intrinsic Dimensionality. J. Mach. Learn. Res. 21 1–38
2020
Later among the works it cites.
[author] Ning, YangY., Sida, PengP. and Imai, KosukeK. (2020). Robust estimation of causal effects via a high-dimensional covariate balancing propensity score. Biometrika 107 533–554
2020
Later among the works it cites.
[author] Rotnitzky, AndreaA. and Smucler, EzequielE. (2020). Efficient Adjustment Sets for Population Average Causal Treatment Effect Estimation in Graphical Models. Journal of machine learning research 21
2020
Later among the works it cites.
[author] Shi, ChengchunC., Lu, WenbinW. and Song, RuiR. (2020). Breaking the Curse of Nonregularity with Subagging - Inference of the Mean Outcome under Optimal Treatment Regimes. Journal of machine learning research 21
2020
Later among the works it cites.
Sondhi, A
2020
Later among the works it cites.
[author] Tang, ZiyangZ., Feng, YihaoY., Li, LihongL., Zhou, DengyongD. and Liu, QiangQ. (2020). Doubly Robust Bias Reduction in Infinite Horizon Off-Policy Estimation. ICLR 2020
2020
Later among the works it cites.
[author] Tennenholtz, GuyG., Shalit, UriU. and Mannor, ShieS. (2020). Off-Policy Evaluation in Partially Observable Environments. Proceedings of the AAAI Conference on Artificial Intelligence 34 10276-10283
2020
Later among the works it cites.
[author] Uehara, MasatoshiM., Huang, JiaweiJ. and Jiang, NanN. (2020). Minimax Weight and Q-Function Learning for Off-Policy Evaluation. ICML 2020
2020
Later among the works it cites.
[author] Ueno, TsuyoshiT., Kawanabe, MotoakiM., Mori, TakeshiT., Maeda, Shin-IchiS.-I. and Ishii, ShinS. (2011). Generalized TD learning. Journal of Machine Learning Research 12 1977–2020
2020
Later among the works it cites.
[author] Wang, YixinY. and Zubizarreta, Jose RJ. R. (2020). Minimal dispersion approximately balancing weights: asymptotic properties and practical considerations. Biometrika 107 93–105
2020
Later among the works it cites.
[author] Yin, MingM. and Wang, Yu-XiangY.-X. (2020). Asymptotically Efficient Off-Policy Evaluation for Tabular Reinforcement Learning. In Proceedings of the 23nd International Workshop on Artificial Intelligence and Statistics
2020
Later among the works it cites.
[author] Bennett, AndrewA. and Kallus, NathanN. (2021). Proximal Reinforcement Learning: Efficient Off-Policy Evaluation in Partially Observed Markov Decision Processes
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Kuzborskij, I
2021
Later among the works it cites.
[author] Liao, PengP., Klasnja, PredragP. and Murphy, SusanS. (2021). Off-policy estimation of long-term average outcomes with applications to mobile health. Journal of the American Statistical Association 116 382–391
2021
Later among the works it cites.
[author] Mo, WeibinW., Qi, ZhenglingZ. and Liu, YufengY. (2021). Learning optimal distributionally robust individualized treatment rules. Journal of the American Statistical Association 116 659–674
2021
Later among the works it cites.
2021
Later among the works it cites.
[author] Pananjady, AshwinA. and Wainwright, Martin JM. J. (2021). Instance-Dependent l ∞ l_{\infty} -Bounds for Policy Evaluation in Tabular Reinforcement Learning. IEEE transactions on information theory 67 566–585
2021
Later among the works it cites.
2021
Later among the works it cites.
[author] Singh, RahulR. (2021). Debiased kernel methods. arXiv preprint arXiv:2102.11076
2021
Later among the works it cites.
2021
Later among the works it cites.
[author] Uehara, MasatoshiM., Imaizumi, MasaakiM., Jiang, NanN., Kallus, NathanN., Sun, WenW. and Xie, TengyangT. (2021). Finite Sample Analysis of Minimax Offline Reinforcement Learning: Completeness, Fast Rates and First-Order Efficiency
2021
Later among the works it cites.
[author] Wu, YunanY. and Wang, LanL. (2021). Resampling-based confidence intervals for model-free robust inference on optimal treatment regimes. Biometrics 77 465–476
2021
Later among the works it cites.
2021
Later among the works it cites.
[author] Zhang, J.J. and Bareinboim, E.E. (2021). Non-Parametric Methods for Partial Identification of Causal Effects. Columbia CausalAI Laboratory Technical Report (R-72)
2021
Later among the works it cites.
2021
Later among the works it cites.
2022
Closest in time.
2022
Closest in time.
[author] Huang, AudreyA. and Jiang, NanN. (2022). Beyond the Return: Off-policy Function Estimation under User-specified Error-measuring Distributions. Neurips
2022
Closest in time.
2022
Closest in time.
[author] Liao, PengP., Qi, ZhenglingZ., Wan, RunzheR., Klasnja, PredragP. and Murphy, SusanS. (2022). Batch policy learning in average reward Markov decision processes. Annals of Statistics accepted
2022
Closest in time.
2022
Closest in time.
Miyaguchi, K
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
[author] Smucler, EzequielE., Sapienza, FacundoF. and Rotnitzky, AndreaA. (2022). Efficient adjustment sets in causal graphical models with hidden variables. Biometrika 109 49–65
2022
Closest in time.
2022
Closest in time.