Fetching the paper…
Reading the bibliography…
The statistical framework of Generalized Linear Models (GLM) can be applied to sequential problems involving categorical or ordinal rewards associated, for instance, with clicks, likes or ratings.
Hedging the drift: Learning to optimize under non-stationarity
W. C. Cheung, D. Simchi-Levi, and R. Zhu · 1903
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
V. Dani, T. P. Hayes, and S. M. Kakade · 2008
Earlier work this paper cites.
Parametric bandits: The generalized linear case
S. Filippi, O. Cappe, A. Garivier, and C. Szepesvári · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Linearly parameterized bandits
P. Rusmevichientong and J. N. Tsitsiklis · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
A. Garivier and E. Moulines · 2011
Earlier work this paper cites.
Stochastic multi-armed-bandit problem with non-stationary rewards
O. Besbes, Y. Gur, and A. Zeevi · 2014
Cited alongside, same era.
Linear thompson sampling revisited
M. Abeille, A. Lazaric, et al · 2017
Cited alongside, same era.
Provably optimal algorithms for generalized linear contextual bandits
L. Li, Y. Lu, and D. Zhou · 2017
Cited alongside, same era.
Adaptively tracking the best arm with an unknown number of distribution changes
P. Auer, P. Gajane, and R. Ortner · 2018
Cited alongside, same era.
A change-detection based framework for piecewise-stationary multi-armed bandit problem
F. Liu, J. Lee, and N. Shroff · 2018
Cited alongside, same era.
Learning contextual bandits in a non-stationary environment
Q. Wu, N. Iyer, and H. Wang · 2018
Cited alongside, same era.
Nearly optimal adaptive procedure with change detection for piecewise-stationary bandit
Y. Cao, Z. Wen, B. Kveton, and Y. Xie · 2019
Later among the works it cites.
Learning to optimize under non-stationarity
W. C. Cheung, D. Simchi-Levi, and R. Zhu · 2019
Later among the works it cites.
On the performance of thompson sampling on logistic bandits
S. Dong, T. Ma, and B. Van Roy · 2019
Later among the works it cites.
Weighted linear bandits for non-stationary environments
Y. Russac, C. Vernade, and O. Cappé · 2019
Later among the works it cites.
Randomized exploration in generalized linear bandits
B. Kveton, M. Zaheer, C. Szepesvari, L. Li, M. Ghavamzadeh, and C. Boutilier · 2020
Closest in time.
A simple approach for non-stationary linear bandits
P. Zhao, L. Zhang, Y. Jiang, and Z.-H. Zhou · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Besson and E. Kaufmann · 2019
Cited alongside, same era.
Closest in time.