Fetching the paper…
Reading the bibliography…
We study the non-stationary stochastic multi-armed bandit problem, where the reward statistics of each arm may change several times during the course of learning.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Adapting to a changing environment: the brownian restless bandits
Aleksandrs Slivkins and Eli Upfal · 2008
Earlier work this paper cites.
Ucb revisited: Improved regret bounds for the stochastic multi-armed bandit problem
Peter Auer and Ronald Ortner · 2010
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
Aurélien Garivier and Eric Moulines · 2011
Earlier work this paper cites.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Cited alongside, same era.
Regret bounds for restless markov bandits
Ronald Ortner, Daniil Ryabko, Peter Auer, and Rémi Munos · 2014
Cited alongside, same era.
Non-stationary reinforcement learning without prior knowledge: an optimal black-box approach
Chen-Yu Wei and Haipeng Luo · 2014
Cited alongside, same era.
Achieving optimal dynamic regret for non-stationary bandits without prior information
Peter Auer, Yifang Chen, Pratik Gajane, Chung-Wei Lee, Haipeng Luo, Ronald Ortner, and Chen-Yu Wei
Cited in the paper.
Adaptively tracking the best bandit arm with an unknown number of distribution changes
Peter Auer, Pratik Gajane, and Ronald Ortner
Cited in the paper.
The non-stationary stochastic multi-armed bandit problem
Robin Allesiardo, Raphaël Féraud, and Odalric-Ambrym Maillard · 2017
Later among the works it cites.
Adaptively tracking the best arm with an unknown number of distribution changes
Peter Auer, Pratik Gajane, and Ronald Ortner · 2018
Later among the works it cites.
A new algorithm for non-stationary contextual bandits: Efficient, optimal and parameter-free
Yifang Chen, Chung-Wei Lee, Haipeng Luo, and Chen-Yu Wei · 2019
Later among the works it cites.
Tracking most severe arm changes in bandits
Joe Suk and Samory Kpotufe · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…