Fetching the paper…
Reading the bibliography…
Counterfactual reasoning from logged data has become increasingly important for many applications such as web advertising or healthcare.
A generalization of sampling without replacement from a finite universe
D. G. Horvitz and D. J. Thompson · 1952
Earlier work this paper cites.
Drug dosage in laboratory animals: a handbook
C. D. Barnes and L. G. Eltherington · 1966
Earlier work this paper cites.
Monotone operators and the proximal point algorithm
R. T. Rockafellar · 1976
Earlier work this paper cites.
A generalized proximal point algorithm for certain non-convex minimization problems
M. Fukushima and H. Mine · 1981
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
D. C. Liu and J. Nocedal · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Semiparametric efficiency in multivariate regression models with missing data
J. Robins and A. Rotnitzky · 1995
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Using the nyström method to speed up kernel machines
C. K. Williams and M. Seeger · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
B. Schölkopf and A. Smola · 2002
Earlier work this paper cites.
The propensity score with continuous treatments
K. Hirano and G. W. Imbens · 2004
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
J. Langford and T. Zhang · 2008
Cited alongside, same era.
New England Journal of Medicine , 360(8):753–764, 2009
Estimation of the Warfarin dose with clinical and pharmacogenetic data · 2009
Cited alongside, same era.
Empirical bernstein bounds and sample variance penalization
A. Maurer and M. Pontil · 2009
Cited alongside, same era.
Doubly robust policy evaluation and learning
M. Dudik, J. Langford, and L. Li · 2011
Cited alongside, same era.
Contextual gaussian process bandit optimization
A. Krause and C. S. Ong · 2011
Cited alongside, same era.
An unbiased offline evaluation of contextual bandit algorithms with generalized linear models
L. Li, W. Chu, J. Langford, T. Moon, and X. Wang · 2012
Cited alongside, same era.
A comparative study of counterfactual estimators
T. Nedelec, N. L. Roux, and V. Perchet · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Y.-X. Wang, A. Agarwal, and M. Dudík · 2017
Later among the works it cites.
Optimization over continuous and multi-dimensional decisions with observational data
D. Bertsimas and C. McCord · 2018
Later among the works it cites.
Deep learning with logged bandit feedback
T. Joachims, A. Swaminathan, and M. de Rijke · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Counterfactual reasoning and learning systems: The example of computational advertising
L. Bottou, J. Peters, J. Quiñonero Candela, D. X. Charles, D. M. Chickering, E. Portugaly, D. Ray, P. Simard, and E. Snelson · 2013
Cited alongside, same era.
Monte Carlo theory, methods and examples
A. Owen · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. Schapire · 2014
Cited alongside, same era.
Personalized dose finding using outcome weighted learning
G. Chen, D. Zeng, and M. R. Kosorok · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Cited alongside, same era.
Large-scale validation of counterfactual learning methods: A test-bed
D. Lefortier, A. Swaminathan, X. Gu, T. Joachims, and M. de Rijke · 2016
Cited alongside, same era.
N. Kallus and A. Zhou · 2018
Later among the works it cites.
Catalyst for gradient-based nonconvex optimization
C. Paquette, H. Lin, D. Drusvyatskiy, J. Mairal, and Z. Harchaoui · 2018
Later among the works it cites.
Understanding the impact of entropy on policy optimization
Z. Ahmed, N. Le Roux, M. Norouzi, and D. Schuurmans · 2019
Later among the works it cites.
Semi-parametric efficient policy learning with continuous actions
M. Demirer, V. Syrgkanis, G. Lewis, and V. Chernozhukov · 2019
Later among the works it cites.
Cab: Continuous adaptive blending for policy evaluation and learning
Y. Su, L. Wang, M. Santacatterina, and T. Joachims · 2019
Later among the works it cites.
Efficient contextual bandits with continuous actions
M. Majzoubi, C. Zhang, R. Chari, A. Krishnamurthy, J. Langford, and A. Slivkins · 2020
Closest in time.
Subgaussian and differentiable importance sampling for off-policy evaluation and learning
A. M. Metelli, A. Russo, and M. Restelli · 2021
Closest in time.