Fetching the paper…
Reading the bibliography…
Information-theoretic Bayesian regret bounds of Russo and Van Roy capture the dependence of regret on prior uncertainty.
William R Thompson · 1933
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
An empirical evaluation of Thompson sampling
Olivier Chapelle and Lihong Li · 2011
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Earlier work this paper cites.
Learning to optimize via information-directed sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
An information-theoretic analysis of Thompson sampling
Daniel Russo and Benjamin Van Roy · 2016
Cited alongside, same era.
Near-optimal regret bounds for Thompson sampling
Shipra Agrawal and Navin Goyal · 2017
Cited alongside, same era.
Provably optimal algorithms for generalized linear contextual bandits
Lihong Li, Yu Lu, and Dengyong Zhou · 2017
Later among the works it cites.
Satisficing in time-sensitive bandit learning
Daniel Russo and Benjamin Van Roy · 2018
Closest in time.
A tutorial on Thompson sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…