Fetching the paper…
Reading the bibliography…
Contextual bandits serve as a fundamental model for many sequential decision making tasks.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Sequential generalized likelihood ratios and adaptive treatment allocation for optimal sequential selection
H. P. Chan and T. L. Lai · 2006
Earlier work this paper cites.
The epoch-greedy algorithm for contextual multi-armed bandits
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
Online linear optimization and adaptive routing
Baruch Awerbuch and Robert Kleinberg · 2008
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P. Hayes, and Sham M. Kakade · 2008
Earlier work this paper cites.
Introduction to Nonparametric Estimation
Alexandre B. Tsybakov · 2008
Earlier work this paper cites.
Online models for content optimization
Deepak Agarwal, Bee-Chung Chen, Pradheep Elango, Nitin Motgi, Seung-Taek Park, Raghu Ramakrishnan, Scott Roy, and Joe Zachariah · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Earlier work this paper cites.
Linearly parameterized bandits
Paat Rusmevichientong and John N Tsitsiklis · 2010
Cited alongside, same era.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
CVXPY: A Python-embedded modeling language for convex optimization
Steven Diamond and Stephen Boyd · 2016
Cited alongside, same era.
Mostly exploration-free algorithms for contextual bandits
Hamsa Bastani, Mohsen Bayati, and Khashayar Khosravi · 2017
Later among the works it cites.
Minimal exploration in structured stochastic bandits
Richard Combes, Stefan Magureanu, and Alexandre Proutiere · 2017
Later among the works it cites.
The end of optimism? an asymptotic analysis of finite-armed linear bandits
Tor Lattimore and Csaba Szepesvári · 2017
Later among the works it cites.
A smoothed analysis of the greedy algorithm for the linear contextual bandit problem
Sampath Kannan, Jamie H Morgenstern, Aaron Roth, Bo Waggoner, and Zhiwei Steven Wu · 2018
Later among the works it cites.
Exploration in structured reinforcement learning
Jungseul Ok, Alexandre Proutiere, and Damianos Tranos · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimal best arm identification with fixed confidence
A. Garivier and E. Kaufmann · 2016
Cited alongside, same era.
Rémy Degenne, Wouter M Koolen, and Pierre Ménard · 2019
Closest in time.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2019
Closest in time.