Fetching the paper…
Reading the bibliography…
This paper studies the stochastic linear bandit problem, where a decision-maker chooses actions from possibly time-dependent sets of vectors in $\mathbb{R}^d$ and receives noisy rewards.
Meta Dynamic Pricing: Transfer Learning Across Experiments arXiv:1902.10918
Bastani, Hamsa, David Simchi-Levi, Ruihao Zhu. 2019 · 1902
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, William R. 1933 · 1933
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, Tze Leung, Herbert Robbins. 1985 · 1985
Earlier work this paper cites.
Introduction to Algorithms
Cormen, Thomas H., Charles E. Leiserson, Ronald L. Rivest, Clifford Stein. 2001 · 2001
Earlier work this paper cites.
A general theory of the stochastic linear bandit and its applications
Hamidi, Nima, Mohsen Bayati. 2020 · 2002
Earlier work this paper cites.
Mots: Minimax optimal thompson sampling
Jin, Tianyuan, Pan Xu, Jieming Shi, Xiaokui Xiao, Quanquan Gu. 2020 · 2003
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, Varsha, Thomas P. Hayes, Sham M. Kakade. 2008 · 2008
Earlier work this paper cites.
Linearly parameterized bandits
Rusmevichientong, Paat, John N Tsitsiklis. 2010 · 2010
Earlier work this paper cites.
A modern bayesian look at the multi-armed bandit
Scott, Steven L. 2010 · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Yasin, Dávid Pál, Csaba Szepesvári. 2011 · 2011
Cited alongside, same era.
Analysis of thompson sampling for the multi-armed bandit problem
Agrawal, Shipra, Navin Goyal. 2012 · 2012
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, Sébastien, Nicolo Cesa-Bianchi, et al. 2012 · 2012
Cited alongside, same era.
Prior-free and prior-dependent regret bounds for thompson sampling
Bubeck, Sébastien, Che-Yu Liu. 2013 · 2013
Cited alongside, same era.
Learning to optimize via posterior sampling
Russo, Daniel, Benjamin Van Roy. 2014 · 2014
Cited alongside, same era.
Peeking at a/b tests: Why it matters, and what to do about it
Johari, Ramesh, Pete Koomen, Leonid Pekelis, David Walsh. 2017 · 2017
Later among the works it cites.
Why adaptively collected data have negative bias and how to correct for it
Nie, Xinkun, Xiaoying Tian, Jonathan Taylor, James Zou. 2018 · 2018
Later among the works it cites.
A tutorial on thompson sampling
Russo, Daniel J., Benjamin Van Roy, Abbas Kazerouni, Ian Osband, Zheng Wen. 2018 · 2018
Later among the works it cites.
High-dimensional probability: An introduction with applications in data science
Vershynin, Roman. 2018 · 2018
Later among the works it cites.
Bandit Algorithms
Lattimore, Tor, Csaba Szepesvari. 2019 · 2019
Later among the works it cites.
Thompson sampling and approximate inference
Phan, My, Yasin Abbasi Yadkori, Justin Domke. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-armed bandit experiments in the online service economy
Scott, Steven L. 2015 · 2015
Cited alongside, same era.
Linear thompson sampling revisited
Abeille, Marc, Alessandro Lazaric, et al. 2017 · 2017
Cited alongside, same era.
Further optimal regret bounds for thompson sampling
Agrawal, Shipra, Navin Goyal. 2013a
Cited in the paper.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, Shipra, Navin Goyal. 2013b
Cited in the paper.
Introduction to multi-armed bandits
Slivkins, Aleksandrs. 2019 · 2019
Later among the works it cites.