Fetching the paper…
Reading the bibliography…
We study the problem of policy evaluation and learning from batched contextual bandit data when treatments are continuous, going beyond previous work on discrete treatments.
A generalization of sampling without replacement froma finite universe
Daniel Horvitz and Donovan Thompson · 1952
Earlier work this paper cites.
Estimating causal eeffect of treatments in randomized and nonrandomized studies
Donald Rubin · 1974
Earlier work this paper cites.
Weak and strong uniform consistency of the kernel estimate of a density and its derivatives
Bernard Silverman · 1978
Earlier work this paper cites.
Probability in Banach Spaces: isoperimetry and processes
Michel Ledoux and Michel Talagrand · 1991
Earlier work this paper cites.
Nonparametric Econometrics
Adrian Pagan and Aman Ullah · 1999
Earlier work this paper cites.
On the global convergence of bfgs method for nonconvex unconstrained optimization problems
Dong-Hui Li and Masao Fukushima · 2000
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
The Propensity Score with Continuous Treatments, in Applied Bayesian Modeling and Causal Inference from Incomplete-Data Perspectives: An Essential Journey with Donald Rubin’s Statistical Family
Keisuke Hirano and Guido Imbens · 2004
Earlier work this paper cites.
The offset tree for learning with partial labels
Alina Beygelzimer and John Langford · 2009
Cited alongside, same era.
Lecture notes on nonparametrics
Bruce Hansen · 2009
Cited alongside, same era.
Estimation of the warfarin dose with clnical and pharmacogenetic data
T E International Warfarin Pharmacogenetics Consortium, Klein, R B Altman, N Eriksson, B F Gage, S E Kimmel, M-T M Lee, N A Limdi, D Page, D M Roden, M J Wagner, M D Caldwell, and Johnson J A · 2009
Cited alongside, same era.
On the complexity of linear predicttion: Risk bounds, margin bounds, and regularization
Sham Kakade, Karthik Sridharan, and Ambuj Tewari · 2009
Cited alongside, same era.
On the convergence of the concave-convex procedure
Gert R. Lanckriet and Bharath K. Sriperumbudur · 2009
Cited alongside, same era.
Comparison of data-driven bandwidth selectors
Byeong Park and J.S. Marron · 2009
Cited alongside, same era.
Doubly robust policy evaluation and optimization
Miroslav Dudik, Dumitru Erhan, John Langford, and Lihong Li · 2014
Later among the works it cites.
Online decision-making with high-dimensional covariates
Hamsa Bastani and Mohsen Bayati · 2015
Later among the works it cites.
Evaluation of the effect of a continuous treatment: A machine learning approach with an application to treatment for traumatic brain injury
Noemi Kreif, Richard Grieve, Ivan Dia, and David Harrison · 2015
Later among the works it cites.
Counterfactual risk minimization
Adith Swaminathan and Thorsten Joachims · 2015
Later among the works it cites.
The self-normalized estimator for counterfactual learning
Adith Swaminathan and Thorsten Joachims · 2015
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011 accf/aha/hrs focused updates incorporated into the acc/aha/esc 2006 guidelines for the management of patients with atrial fibrillation
Valentin Fuster, Lars E. Ryden, Davis S. Cannom, Harry J. Crijns, Anne B. Curtis, Kenneth A. Ellenbogen, Jonathan L. Halperin, G. Neal Kay, Jean-Yves Le Huezey, James E. Lowe, S. Bertil Olsson, Eric N. Prystowsky, Juan Luis Tamargo, and L. Samuel Wann · 2011
Cited alongside, same era.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Cited alongside, same era.
Later among the works it cites.
Recursive partitioning for personalization using observation data
Nathan Kallus · 2017
Later among the works it cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudik · 2017
Later among the works it cites.