Fetching the paper…
Reading the bibliography…
Contextual bandit algorithms are sensitive to the estimation method of the outcome model as well as the exploration method used, particularly in the presence of rich heterogeneity or complex outcome models, which can lead to difficult estimation problems along the path of learning.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. Thompson · 1933
Earlier work this paper cites.
Handbook of mathematical functions: with formulas, graphs, and mathematical tables
Milton Abramowitz and Irene A Stegun · 1964
Earlier work this paper cites.
Ridge regression: Applications to nonorthogonal problems
A. Hoerl and R. Kennard · 1970
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. Lai and H. Robbins · 1985
Earlier work this paper cites.
Matrix analysis
Roger A Horn, Roger A Horn, and Charles R Johnson · 1990
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
Adjusting for nonignorable drop-out using semiparametric nonresponse models
Daniel O Scharfstein, Andrea Rotnitzky, and James M Robins · 1999
Earlier work this paper cites.
Random forests
L. Breiman · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
P. Auer · 2003
Earlier work this paper cites.
Learning and evaluating classifiers under sample selection bias
B. Zadrozny · 2004
Earlier work this paper cites.
Correcting sample selection bias by unlabeled data
J. Huang, A. Gretton, K. M. Borgwardt, B. Scholkopf, and A. J. Smola · 2007
Earlier work this paper cites.
The bayesian lasso
Trevor Park and George Casella · 2008
Earlier work this paper cites.
Dealing with limited overlap in estimation of average treatment effects
Richard K Crump, V Joseph Hotz, Guido W Imbens, and Oscar A Mitnik · 2009
Earlier work this paper cites.
Learning bounds for importance eeighting
C. Cortes, Y. Mansour, and M. Mohri · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. Schapire · 2010
Earlier work this paper cites.
Nonparametric bandits with covariates
P. Rigollet and R. Zeevi · 2010
Earlier work this paper cites.
A modern bayesian look at the multi-armed bandit
S. Scott · 2010
Earlier work this paper cites.
Learning from logged implicit exploration data
A. Strehl, J. Langford, L. Li, and S. Kakade · 2010
Earlier work this paper cites.
Learning from logged implicit exploration data
Alex Strehl, John Langford, Lihong Li, and Sham M Kakade · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
An empirical evaluation of thompson sampling
O. Chapelle and L. Li · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Cited alongside, same era.
Doubly robust policy evaluation and learning
M. Dudik, J. Langford, and L. Li · 2011
Cited alongside, same era.
Analysis of thompson sampling for the multi-armed bandit problem
S. Agrawal and N. Goyal · 2012
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Cited alongside, same era.
Recursive partitioning for heterogeneous causal effects
S. Athey and G. Imbens · 2016
Later among the works it cites.
Random forest for the contextual bandit problem
R. Feraud, R. Allesiardo, T. Urvoy, and F. Clerot · 2016
Later among the works it cites.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Later among the works it cites.
Causal bandits: Learning good interventions via causal inference
F. Lattimore, T. Lattimore, and M. D. Reid · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
P. Thomas and E. Brunskill · 2016
Later among the works it cites.
Making contextual decisions with low technical debt
A. Agarwal, S. Bird, M. Cozowicz, L. Hoang, J. Langford, S. Lee, J. Li, D. Melamed, G. Oshri, O. Ribas, S. Sen, and A. Slivkins · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An unbiased offline evaluation of contextual bandit algorithms with generalized linear models
L. Li, W. Chu, J. Langford, T. Moon, and X. Wang · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawal and N. Goyal · 2013
Cited alongside, same era.
A linear response bandit problem
A. Goldenshluger and A. Zeevi · 2013
Cited alongside, same era.
The multi-armed bandit problem with covariates
V. Perchet and P. Rigollet · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. Schapire · 2014
Cited alongside, same era.
Doubly robust policy evaluation and optimization
M. Dudik, D. Erhan, J. Langford, and L. Li · 2014
Cited alongside, same era.
Approximate residual balancing
S. Athey, G. Imbens, and S. Wager · 2017
Closest in time.
Generalized random forests
S. Athey, J. Tibshirani, and S. Wager · 2017
Closest in time.
Efficient policy learning
S. Athey and S. Wager · 2017
Closest in time.
Accurate inference in adaptive linear models
Y. Deshpande, L. Mackey, V. Syrgkanis, and M. Taddy · 2017
Closest in time.
A practical method for solving contextual bandit problems using decision trees
A. Elmachtoub, R. McNellis, S. Oh, and M. Petrik · 2017
Closest in time.
Counterfactual data-fusion for online reinforcement learners
A. Forney, J. Pearl, and E. Bareinboim · 2017
Closest in time.
Balanced policy evaluation and learning
N. Kallus · 2017
Closest in time.
An actor-critic contextual bandit algorithm for personalized mobile health interventions
H. Lei, A. Tewari, and S. Murphy · 2017
Closest in time.
Provably optimal algorithms for generalized linear contextual bandits
L. Li, Y. Lu, and D. Zhou · 2017
Closest in time.
A tutorial on Thompson sampling
D. Russo, B. Van Roy, A. Kazerouni, I. Osband, and Z. Wen · 2017
Closest in time.
Optimal and adaptive off-policy evaluation in contextual bandits
Y. X. Wang, A. Agarwal, and M. Dudik · 2017
Closest in time.
Alberto Bietti, Alekh Agarwal, and John Langford · 2018
Closest in time.
Policy evaluation and optimization with continuous treatments
Nathan Kallus and Angela Zhou · 2018
Closest in time.
Why adaptively collected data have negative bias and how to correct for it
X. Nie, X. Tian, J. Taylor, and J. Zou · 2018
Closest in time.
A tutorial on thompson sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, Zheng Wen, et al · 2018
Closest in time.
Offline multi-action policy learning: Generalization and optimization
Zhengyuan Zhou, Susan Athey, and Stefan Wager · 2018
Closest in time.