Fetching the paper…
Reading the bibliography…
Contextual bandit algorithms are essential for solving many real-world interactive machine learning problems.
W. R. Thompson · 1933
Earlier work this paper cites.
Beating the hold-out: Bounds for k-fold and progressive cross-validation
A. Blum, A. Kalai, and J. Langford · 1999
Earlier work this paper cites.
Online bagging and boosting
N. C. Oza and S. Russell · 2001
Earlier work this paper cites.
Risk bounds for statistical learning
P. Massart, É. Nédélec, et al · 2006
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
J. Langford and T. Zhang · 2008
Earlier work this paper cites.
On the generalization ability of online strongly convex programming algorithms
S. M. Kakade and A. Tewari · 2009
Earlier work this paper cites.
Algorithms for active learning
D. J. Hsu · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
An empirical evaluation of thompson sampling
O. Chapelle and L. Li · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Online importance weight aware updates
N. Karampatziakis and J. Langford · 2011
Earlier work this paper cites.
Contextual bandit learning with predictable rewards
A. Agarwal, M. Dudík, S. Kale, J. Langford, and R. E. Schapire · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawal and N. Goyal · 2013
Cited alongside, same era.
Ad click prediction: a view from the trenches
H. B. McMahan, G. Holt, D. Sculley, M. Young, D. Ebner, J. Grady, L. Nie, T. Phillips, E. Davydov, D. Golovin, et al · 2013
Cited alongside, same era.
Efficient online bootstrapping for large scale learning
Z. Qin, V. Petricek, N. Karampatziakis, L. Li, and J. Langford · 2013
Cited alongside, same era.
Normalized online learning
S. Ross, P. Mineiro, and J. Langford · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. E. Schapire · 2014
The self-normalized estimator for counterfactual learning
A. Swaminathan and T. Joachims · 2015
Later among the works it cites.
A multiworld testing decision service
A. Agarwal, S. Bird, M. Cozowicz, L. Hoang, J. Langford, S. Lee, J. Li, D. Melamed, G. Oshri, O. Ribas, et al · 2016
Later among the works it cites.
Recommendations as treatments: debiasing learning and evaluation
T. Schnabel, A. Swaminathan, A. Singh, N. Chandak, and T. Joachims · 2016
Later among the works it cites.
Open problem: First-order regret bounds for contextual bandits
A. Agarwal, A. Krishnamurthy, J. Langford, H. Luo, et al · 2017
Later among the works it cites.
Make the minority great again: First-order regret bound for contextual bandits
Z. Allen-Zhu, S. Bubeck, and Y. Li · 2018
Closest in time.
Practical contextual bandits with regression oracles
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Thompson sampling with the online bootstrap
D. Eckles and M. Kaptein · 2014
Cited alongside, same era.
Theory of disagreement-based active learning
S. Hanneke · 2014
Cited alongside, same era.
Practical lessons from predicting clicks on ads at facebook
X. He, J. Pan, O. Jin, T. Xu, B. Liu, T. Xu, Y. Shi, A. Atallah, R. Herbrich, S. Bowers, et al · 2014
Cited alongside, same era.
Efficient and parsimonious agnostic active learning
T.-K. Huang, A. Agarwal, D. J. Hsu, J. Langford, and R. E. Schapire · 2015
Cited alongside, same era.
Next: A system for real-world development, evaluation, and application of active learning
K. G. Jamieson, L. Jain, C. Fernandez, N. J. Glattard, and R. Nowak · 2015
Cited alongside, same era.
Bootstrapped thompson sampling and deep exploration
I. Osband and B. Van Roy · 2015
Cited alongside, same era.
D. J. Foster, A. Agarwal, M. Dudík, H. Luo, and R. E. Schapire · 2018
Closest in time.
A smoothed analysis of the greedy algorithm for the linear contextual bandit problem
S. Kannan, J. Morgenstern, A. Roth, B. Waggoner, and Z. S. Wu · 2018
Closest in time.
A tutorial on thompson sampling
D. J. Russo, B. Van Roy, A. Kazerouni, I. Osband, Z. Wen, et al · 2018
Closest in time.
Bootstrap thompson sampling and sequential decision problems in the behavioral sciences
D. Eckles and M. Kaptein · 2019
Closest in time.
Active learning for cost-sensitive classification
A. Krishnamurthy, A. Agarwal, T.-K. Huang, H. Daume III, and J. Langford · 2019
Closest in time.
Mostly exploration-free algorithms for contextual bandits
H. Bastani, M. Bayati, and K. Khosravi · 2021
Closest in time.