Fetching the paper…
Reading the bibliography…
Recent growing adoption of experimentation in practice has led to a surge of attention to multiarmed bandits as a technique to reduce the opportunity cost of online experiments.
Online decision making with high-dimensional covariates
Bastani, Hamsa, Mohsen Bayati. 2020 · 1902
Earlier work this paper cites.
Adaptive exploration in linear contextual bandit
Hao, Botao, Tor Lattimore, Csaba Szepesvari. 2019 · 1910
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, William R. 1933 · 1933
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, Tze Leung, Herbert Robbins. 1985 · 1985
Earlier work this paper cites.
Associative reinforcement learning using linear probabilistic concepts
Abe, Naoki, Philip M. Long. 1999 · 1999
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Laurent, Beatrice, Pascal Massart. 2000 · 2000
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, Peter. 2003 · 2003
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, Varsha, Thomas P. Hayes, Sham M. Kakade. 2008 · 2008
Earlier work this paper cites.
Woodroofe’s one-armed bandit problem revisited
Goldenshluger, Alexander, Assaf Zeevi. 2009 · 2009
Earlier work this paper cites.
Linearly parameterized bandits
Rusmevichientong, Paat, John N Tsitsiklis. 2010 · 2010
Earlier work this paper cites.
A modern bayesian look at the multi-armed bandit
Scott, Steven L. 2010 · 2010
Cited alongside, same era.
Gaussian process optimization in the bandit setting: No regret and experimental design
Srinivas, Niranjan, Andreas Krause, Sham Kakade, Matthias Seeger. 2010 · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Yasin, Dávid Pál, Csaba Szepesvári. 2011 · 2011
Cited alongside, same era.
Online Learning for Linearly Parametrized Control Problems
Abbasi-Yadkori, Yasin. 2012 · 2012
Cited alongside, same era.
User-friendly tail bounds for sums of random matrices
Tropp, Joel A. 2012 · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, Shipra, Navin Goyal. 2013 · 2013
Cited alongside, same era.
An information-theoretic analysis of thompson sampling
Russo, Daniel, Benjamin Van Roy. 2016 · 2016
Later among the works it cites.
Linear thompson sampling revisited
Abeille, Marc, Alessandro Lazaric, et al. 2017 · 2017
Later among the works it cites.
Mostly exploration-free algorithms for contextual bandits
Bastani, Hamsa, Mohsen Bayati, Khashayar Khosravi. 2017 · 2017
Later among the works it cites.
Peeking at a/b tests: Why it matters, and what to do about it
Johari, Ramesh, Pete Koomen, Leonid Pekelis, David Walsh. 2017 · 2017
Later among the works it cites.
An information-theoretic analysis for thompson sampling with many actions
Dong, Shi, Benjamin Van Roy. 2018 · 2018
Later among the works it cites.
A smoothed analysis of the greedy algorithm for the linear contextual bandit problem
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A linear response bandit problem
Goldenshluger, Alexander, Assaf Zeevi. 2013 · 2013
Cited alongside, same era.
Learning to optimize via posterior sampling
Russo, Daniel, Benjamin Van Roy. 2014 · 2014
Cited alongside, same era.
The pareto regret frontier for bandits
Lattimore, Tor. 2015 · 2015
Cited alongside, same era.
Multi-armed bandit experiments in the online service economy
Scott, Steven L. 2015 · 2015
Cited alongside, same era.
Kannan, Sampath, Jamie H Morgenstern, Aaron Roth, Bo Waggoner, Zhiwei Steven Wu. 2018 · 2018
Later among the works it cites.
Information directed sampling and bandits with heteroscedastic noise
Kirschner, Johannes, Andreas Krause. 2018 · 2018
Later among the works it cites.
The externalities of exploration and how data diversity helps exploitation
Raghavan, Manish, Aleksandrs Slivkins, Jennifer Wortman Vaughan, Zhiwei Steven Wu. 2018 · 2018
Later among the works it cites.
The unreasonable effectiveness of greedy algorithms in multi-armed bandit with many arms
Bayati, Mohsen, Nima Hamidi, Ramesh Johari, Khashayar Khosravi. 2020 · 2020
Closest in time.