Fetching the paper…
Reading the bibliography…
In bandit with distribution shifts, one aims to automatically adapt to unknown changes in reward distribution, and restart exploration when necessary.
Introduction to Multi-Armed Bandits
Aleksandrs Slivkins · 1904
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Discounted ucb
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Adapting to a changing environment: the brownian restless bandits
Alex Slivkins and Eli Upfal · 2008
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
Aurélien Garivier and Eric Moulines · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicoló Cesa-Bianchi · 2012
Earlier work this paper cites.
The best of both worlds: Stochastic and adversarial bandits
Sébastien Bubeck and Aleksandrs Slivkins · 2012
Earlier work this paper cites.
One practical algorithm for both stochastic and adversarial bandits
Yevgeny Seldin and Aleksandrs Slivkins · 2014
Earlier work this paper cites.
An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits
Peter Auer and Chao-Kai Chiang · 2016
Cited alongside, same era.
Tight policy regret bounds for improving and decaying bandits
Hoda Heidari, Michael Kearns, and Aaron Roth · 2016
Cited alongside, same era.
The non-stationary stochastic multi-armed bandit problem
Robin Allesiardo, Raphaël Féraud, and Odalric-Ambrym Maillard · 2017
Cited alongside, same era.
Rotting bandits
Nir Levine, Koby Crammer, and Shie Mannor · 2017
Cited alongside, same era.
Adaptively tracking the best arm with an unknown number of distribution changes
Peter Auer, Pratik Gajane, and Ronald Ortner · 2018
Cited alongside, same era.
On abruptly-changing and slowly-varying multiarmed bandit problems
Lai Wei and Vaihbav Srivatsva · 2018
Cited alongside, same era.
Rotting bandits are no harder than stochastic ones
Julien Seznec, Andrea Locatelli, Alexandra Carpentier, Alessandro Lazaric, and Michal Valko · 2019
Later among the works it cites.
Open problem: Model selection for contextual bandits
Dylan J. Foster, Akshay Krishnamurthy, and Haipeng Luo · 2020
Later among the works it cites.
Bandit Algoritms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
A single algorithm for both restless and rested rotting bandits
Julien Seznec, Pierre Menard, Alessandro Lazaric, and Michal Valko · 2020
Later among the works it cites.
On slowly-varying non-stationary bandits
Ramakrishnan Krishnamurthy and Aditya Gopalan · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adaptively tracking the best bandit arm with an unknown number of distribution changes
Peter Auer, Pratik Gajane, and Ronald Ortner · 2019
Cited alongside, same era.
A new algorithm for non-stationary contextual bandits: Efficient, optimal and parameter-free
Yifang Chen, Chung-Wei Lee, Haipeng Luo, and Chen-Yu Wei · 2019
Cited alongside, same era.
Learning to optimize under non-stationarity
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2019
Cited alongside, same era.
Distribution-dependent and time-uniform bounds for piecewise i.i.d bandits
Subhojyoti Mukherjee and Odalric-Ambrym Maillard · 2019
Cited alongside, same era.
Anne Gael Manegueu, Alexandra Carpentier, and Yi Yu · 2021
Closest in time.
The pareto frontier of model selection for general contextual bandits
Teodor Marinov and Julian Zimmert · 2021
Closest in time.
Non-stationary reinforcement learning without prior knowledge: An optimal black-box approach
Chen-Yu Wei and Haipeng Luo · 2021
Closest in time.
A new look at dynamic regret for non-stationary stochastic bandits
Yasin Abbasi-Yadkori, András György, and Nevena Lazic · 2022
Closest in time.
Efficient change-point detection for tackling piecewise-stationary bandits
Lilian Besson, Emilie Kaufmann, Odalric-Ambrym Maillard, and Julien Seznec · 2022
Closest in time.