Fetching the paper…
Reading the bibliography…
Motivated by recommendation problems in music streaming platforms, we propose a nonstationary stochastic bandit model in which the expected reward of an arm depends on the number of rounds that have passed since the arm was last pulled.
Bandit processes and dynamic allocation indices
John C Gittins · 1979
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Restless bandits: Activity allocation in a changing world
Peter Whittle · 1988
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Minimizing service and operation costs of periodic scheduling
Amotz Bar-Noy, Randeep Bhatia, Joseph (Seffi) Naor, and Baruch Schieber · 2002
Earlier work this paper cites.
Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
Eyal Even-Dar, Shie Mannor, and Yishay Mansour · 2006
Earlier work this paper cites.
Learning diverse rankings with multi-armed bandits
Filip Radlinski, Robert Kleinberg, and Thorsten Joachims · 2008
Earlier work this paper cites.
Online bandit learning against an adaptive adversary: from regret to policy regret
Raman Arora, Ofer Dekel, and Ambuj Tewari · 2012
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolò Cesa-Bianchi · 2012
Cited alongside, same era.
Online learning of rested and restless bandits
Cem Tekin and Mingyan Liu · 2012
Cited alongside, same era.
Online learning with switching costs and other adaptive adversaries
Nicolo Cesa-Bianchi, Ofer Dekel, and Ohad Shamir · 2013
Cited alongside, same era.
Multi-armed bandit problem with known trend
Djallel Bouneffouf and Raphael Féraud · 2016
Cited alongside, same era.
Tight policy regret bounds for improving and decaying bandits
Hoda Heidari, Michael Kearns, and Aaron Roth · 2016
Cited alongside, same era.
Discrepancy-based algorithms for non-stationary rested bandits
Corinna Cortes, Giulia DeSalvo, Vitaly Kuznetsov, Mehryar Mohri, and Scott Yang · 2017
Later among the works it cites.
Rotting bandits
Nir Levine, Koby Crammer, and Shie Mannor · 2017
Later among the works it cites.
Recharging bandits
Robert Kleinberg and Nicole Immorlica · 2018
Later among the works it cites.
Fighting boredom in recommender systems with linear reinforcement learning
Romain Warlop, Alessandro Lazaric, and Jérémie Mary · 2018
Later among the works it cites.
Blocking bandits
Soumya Basu, Rajat Sen, Sujay Sanghavi, and Sanjay Shakkottai · 2019
Closest in time.
Recovering bandits
Ciara Pike-Burke and Steffen Grunewalder · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
DCM bandits: Learning to rank with multiple clicks
Sumeet Katariya, Branislav Kveton, Csaba Szepesvari, and Zheng Wen · 2016
Cited alongside, same era.
Cascading bandits: Learning to rank in the cascade model
Branislav Kveton, Csaba Szepesvari, Zheng Wen, and Azin Ashkan
Cited in the paper.
Tight regret bounds for stochastic combinatorial semi-bandits
Branislav Kveton, Zheng Wen, Azin Ashkan, and Csaba Szepesvari
Cited in the paper.
Rotting bandits are no harder than stochastic ones
Julien Seznec, Andrea Locatelli, Alexandra Carpentier, Alessandro Lazaric, and Michal Valko · 2019
Closest in time.