Fetching the paper…
Reading the bibliography…
We introduce algorithms that achieve state-of-the-art \emph{dynamic regret} bounds for non-stationary linear stochastic bandit setting.
Using confidence bounds for exploitation-exploration trade-offs
P. Auer · 2002
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. Schapire · 2002
Earlier work this paper cites.
Prediction, Learning, and Games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
V. Dani, T. Hayes, and S. Kakade · 2008
Earlier work this paper cites.
Linearly parameterized bandits
P. Rusmevichientong and J. Tsitsiklis · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
W. Chu, L. Li, L. Reyzin, and R. Schapire · 2011
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
A. Garivier and E. Moulines · 2011
Earlier work this paper cites.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
Online optimization with gradual variations
C. Chiang, T. Yang, C. Lee, M. Mahdavi, C. Lu, R. Jin, and S. Zhu · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawal and N. Goyal · 2013
Cited alongside, same era.
Stochastic multi-armed bandit with non-stationary rewards
O. Besbes, Y. Gur, and A. Zeevi · 2014
Cited alongside, same era.
Learning to optimize via posterior sampling
D. Russo and B. V. Roy · 2014
Cited alongside, same era.
Non-stationary stochastic optimization
O. Besbes, Y. Gur, and A. Zeevi · 2015
Cited alongside, same era.
Tracking the best expert in non-stationary stochastic environments
C.-Y. Wei, Y.-T. Hong, and C.-J. Lu · 2016
Later among the works it cites.
Linear thompson sampling revisited
M. Abeille and A. Lazaric · 2017
Later among the works it cites.
Corralling a band of bandit algorithms
A. Agarwal, H. Luo, B. Neyshabur, and R. E. Schapire · 2017
Later among the works it cites.
Optimal exploration-exploitation in a multi-armed-bandit problem with non-stationary rewards
O. Besbes, Y. Gur, and A. Zeevi · 2018
Closest in time.
Bandit Algorithms
T. Lattimore and C. Szepesvári · 2018
Closest in time.
Efficient contextual bandits in non-stationary worlds
H. Luo, C. Wei, A. Agarwal, and J. Langford · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Online optimization : Competing with dynamic comparators
A. Jadbabaie, A. Rakhlin, S. Shahrampour, and K. Sridharan · 2015
Cited alongside, same era.
Multi-armed bandits: Competing with optimal sequences
Z. Karnin and O. Anava · 2016
Cited alongside, same era.
Chasing demand: Learning and earning in a changing environments
N. Keskin and A. Zeevi · 2016
Cited alongside, same era.
High Dimensional Statistics
R. Rigollet and J. Hütter · 2018
Closest in time.
Regret bounds for generalized linear bandits under parameter drift
L. Faury, Y. Russac, M. Abeille, and C. Calauzenes · 2021
Closest in time.