Fetching the paper…
Reading the bibliography…
We consider off-policy evaluation (OPE) in continuous treatment settings, such as personalized dose-finding.
Van Der Vaart, A. W. & Wellner, J. A. (1996), Weak convergence, in
1996
Earlier work this paper cites.
Massart, P. et al. (2000), ‘About the constants in talagrand’s concentration inequalities for empirical processes’, Annals of Probability
2000
Earlier work this paper cites.
Murphy, S. A., van der Laan, M. J., Robins, J. M. & Group, C. P. P. R. (2001), ‘Marginal mean models for dynamic regimes’, Journal of the American Statistical Association
2001
Earlier work this paper cites.
Hirano, K., Imbens, G. W. & Ridder, G. (2003), ‘Efficient estimation of average treatment effects using the estimated propensity score’, Econometrica
2003
Earlier work this paper cites.
Murphy, S. A. (2003), ‘Optimal dynamic treatment regimes’, Journal of the Royal Statistical Society: Series B (Statistical Methodology)
2003
Earlier work this paper cites.
2004
Earlier work this paper cites.
Adamczak, R. et al. (2008), ‘A tail inequality for suprema of unbounded empirical processes with applications to markov chains’, Electronic Journal of Probability
2008
Earlier work this paper cites.
Friedrich, F., Kempe, A., Liebscher, V. & Winkler, G. (2008), ‘Complexity penalized m-estimation: fast computation’, Journal of Computational and Graphical Statistics
2008
Earlier work this paper cites.
Boysen, L., Kempe, A., Liebscher, V., Munk, A., Wittich, O. et al. (2009), ‘Consistencies and rates of convergence of jump-penalized least squares estimators’, The Annals of Statistics
2009
Earlier work this paper cites.
Consortium, I. W. P. (2009), ‘Estimation of the warfarin dose with clinical and pharmacogenetic data’, New England Journal of Medicine
2009
Earlier work this paper cites.
2009
Earlier work this paper cites.
Harchaoui, Z. & Lévy-Leduc, C. (2010), ‘Multiple change-point estimation with a total variation penalty’, Journal of the American Statistical Association
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
2011
Earlier work this paper cites.
2011
Earlier work this paper cites.
Li, L., Chu, W., Langford, J. & Wang, X. (2011), Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms, in
2011
Earlier work this paper cites.
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M. & Duchesnay, E. (2011), ‘Scikit-learn: Machine learning in Python’, Journal of Machine Learning Research
2011
Earlier work this paper cites.
Qian, M. & Murphy, S. A. (2011), ‘Performance guarantees for individualized treatment rules’, Annals of statistics
2011
Earlier work this paper cites.
Killick, R., Fearnhead, P. & Eckley, I. A. (2012), ‘Optimal detection of changepoints with a linear computational cost’, Journal of the American Statistical Association
2012
Earlier work this paper cites.
Wang, L., Rotnitzky, A., Lin, X., Millikan, R. E. & Thall, P. F. (2012), ‘Evaluation of viable dynamic treatment regimes in a sequentially randomized trial of advanced prostate cancer’, Journal of the American Statistical Association
2012
Earlier work this paper cites.
Zhang, B., Tsiatis, A. A., Laber, E. B. & Davidian, M. (2012), ‘A robust method for estimating optimal treatment regimes’, Biometrics
2012
Earlier work this paper cites.
Chakraborty, B. & Moodie, E. (2013), Statistical methods for dynamic treatment regimes
2013
Cited alongside, same era.
Zhang, B., Tsiatis, A. A., Laber, E. B. & Davidian, M. (2013), ‘Robust estimation of optimal dynamic treatment regimes for sequential treatment decisions’, Biometrika
2013
Cited alongside, same era.
Chernozhukov, V., Chetverikov, D., Kato, K. et al. (2014), ‘Gaussian approximation of suprema of empirical processes’, The Annals of Statistics
2014
Cited alongside, same era.
Dudík, M., Erhan, D., Langford, J., Li, L. et al. (2014), ‘Doubly robust policy evaluation and optimization’, Statistical Science
2014
Cited alongside, same era.
With 32 discussions by 47 authors and a rejoinder by the authors
Frick, K., Munk, A. & Sieling, H. (2014), ‘Multiscale change point inference’, J. R. Stat. Soc. Ser. B. Stat. Methodol · 2014
Cited alongside, same era.
Shi, C., Fan, A., Song, R. & Lu, W. (2018), ‘High-dimensional a-learning for optimal dynamic treatment regimes’, Annals of statistics
2018
Later among the works it cites.
Wang, L., Zhou, Y., Song, R. & Sherwood, B. (2018), ‘Quantile-optimal treatment regimes’, Journal of the American Statistical Association
2018
Later among the works it cites.
2018
Later among the works it cites.
Chernozhukov, V., Demirer, M., Lewis, G. & Syrgkanis, V. (2019), ‘Semi-parametric efficient policy learning with continuous actions’, Advances in Neural Information Processing Systems
2019
Later among the works it cites.
Imaizumi, M. & Fukumizu, K. (2019), Deep neural networks learn non-smooth functions effectively, in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fryzlewicz, P. (2014), ‘Wild binary segmentation for multiple change-point detection’, Ann. Statist
2014
Cited alongside, same era.
Schulte, P. J., Tsiatis, A. A., Laber, E. B. & Davidian, M. (2014), ‘Q-and a-learning methods for estimating optimal dynamic treatment regimes’, Statistical science: a review journal of the Institute of Mathematical Statistics
2014
Cited alongside, same era.
LeCun, Y., Bengio, Y. & Hinton, G. (2015), ‘Deep learning’, nature
2015
Cited alongside, same era.
Chen, G., Zeng, D. & Kosorok, M. R. (2016), ‘Personalized dose finding using outcome weighted learning’, Journal of the American Statistical Association
2016
Cited alongside, same era.
Jiang, N. & Li, L. (2016), Doubly robust off-policy value evaluation for reinforcement learning, in
2016
Cited alongside, same era.
Luedtke, A. R. & Van Der Laan, M. J. (2016), ‘Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy’, Annals of statistics
2016
Cited alongside, same era.
Niu, Y. S., Hao, N. & Zhang, H. (2016), ‘Multiple change-point detection: A selective overview’, Statistical Science
2016
Cited alongside, same era.
2019
Later among the works it cites.
Cai, H., Lu, W. & Song, R. (2020), On validation and planning of an optimal decision rule with application in healthcare studies, in
2020
Closest in time.
den Boer, A. V. & Keskin, N. B. (2020), ‘Discontinuous demand functions: estimation and pricing’, Management Science
2020
Closest in time.
Kallus, N. & Uehara, M. (2020 a
2020
Closest in time.
Kallus, N. & Uehara, M. (2020 b
2020
Closest in time.
Majzoubi, M., Zhang, C., Chari, R., Krishnamurthy, A., Langford, J. & Slivkins, A. (2020), ‘Efficient contextual bandits with continuous actions’, Advances in Neural Information Processing Systems
2020
Closest in time.
Schmidt-Hieber, J. et al. (2020), ‘Nonparametric regression using deep neural networks with relu activation function’, Annals of Statistics
2020
Closest in time.
Shi, C., Lu, W. & Song, R. (2020), ‘Breaking the curse of nonregularity with subagging: inference of the mean outcome under optimal treatment regimes’, Journal of Machine Learning Research
2020
Closest in time.
Sondhi, A., Arbour, D. & Dimmery, D. (2020), Balanced off-policy evaluation in general action spaces, in
2020
Closest in time.
Su, Y., Dimakopoulou, M., Krishnamurthy, A. & Dudík, M. (2020), Doubly robust off-policy evaluation with shrinkage, in
2020
Closest in time.
Su, Y., Srinath, P. & Krishnamurthy, A. (2020), Adaptive estimator selection for off-policy evaluation, in
2020
Closest in time.
Wu, Y. & Wang, L. (2020), ‘Resampling-based confidence intervals for model-free robust inference on optimal treatment regimes’, Biometrics
2020
Closest in time.
Zhu, L., Lu, W., Kosorok, M. R. & Song, R. (2020), Kernel assisted learning for personalized dose finding, in
2020
Closest in time.
Zhu, L., Lu, W. & Song, R. (2020), Causal effect estimation and optimal dose suggestions in mobile health, in
2020
Closest in time.
Farrell, M. H., Liang, T. & Misra, S. (2021), ‘Deep neural networks for estimation and inference’, Econometrica
2021
Closest in time.
Shi, C., Zhang, S., Lu, W. & Song, R. (2021), ‘Statistical inference of the value function for reinforcement learning in infinite-horizon settings’, Journal of the Royal Statistical Society. Series B: Statistical Methodology
2021
Closest in time.
Krishnamurthy, A., Langford, J., Slivkins, A. & Zhang, C. (2019), Contextual bandits with continuous actions: Smoothing, zooming, and adapting, in
2027
Closest in time.