Fetching the paper…
Reading the bibliography…
In stochastic multi-armed bandits, the reward distribution of each arm is assumed to be stationary.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze L. Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Learning from time-changing data with adaptive windowing
Albert Bifet and Ricard Gavaldà · 2007
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
Jean-Yves Audibert and Sébastien Bubeck · 2009
Earlier work this paper cites.
On upper-confidence-bound policies for switching bandit problems
Aurélien Garivier and Eric Moulines · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolò Cesa-Bianchi · 2012
Cited alongside, same era.
Stochastic multi-armed bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Cited alongside, same era.
Multi-armed bandit problem with known trend
Djallel Bouneffouf and Raphael Féraud · 2016
Cited alongside, same era.
Tight policy regret bounds for improving and decaying bandits
Hoda Heidari, Michael Kearns, and Aaron Roth · 2016
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer
Cited in the paper.
The nonstochastic multi-armed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire
Cited in the paper.
Rotting bandits
Nir Levine, Koby Crammer, and Shie Mannor · 2017
Later among the works it cites.
SMPyBandits: An open-source research framework for single and multi-players multi-arm bandit algorithms in Python
Lilian Besson · 2018
Closest in time.
Fighting boredom in recommender systems with linear reinforcement learning
Romain Warlop, Alessandro Lazaric, and Jérémie Mary · 2018
Closest in time.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…