Fetching the paper…
Reading the bibliography…
We consider a stochastic linear bandit model in which the available actions correspond to arbitrary context vectors whose associated rewards follow a non-stationary linear regression model.
Using confidence bounds for exploitation-exploration trade-offs
P. Auer · 2002
Earlier work this paper cites.
Discounted ucb
L. Kocsis and C. Szepesvári · 2006
Earlier work this paper cites.
Self-normalized processes: Limit theory and Statistical Applications
V. H. Peña, T. L. Lai, and Q.-M. Shao · 2008
Earlier work this paper cites.
Piecewise-stationary bandit problems with side observations
J. Y. Yu and S. Mannor · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
The group fused lasso for multiple change-point detection
K. Bleakley and J.-P. Vert · 2011
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
A. Garivier and E. Moulines · 2011
Earlier work this paper cites.
Thompson sampling for dynamic multi-armed bandits
N. Gupta, O.-C. Granmo, and A. Agrawala · 2011
Earlier work this paper cites.
A linear response bandit problem
A. Goldenshluger and A. Zeevi · 2013
Earlier work this paper cites.
Stochastic multi-armed-bandit problem with non-stationary rewards
O. Besbes, Y. Gur, and A. Zeevi · 2014
Cited alongside, same era.
Non-stationary stochastic optimization
O. Besbes, Y. Gur, and A. Zeevi · 2015
Cited alongside, same era.
Attribution modeling increases efficiency of bidding in display advertising
Diemert Eustache, Meynet Julien, P. Galland, and D. Lefortier · 2017
Cited alongside, same era.
Chasing demand: Learning and earning in a changing environment
N. B. Keskin and A. Zeevi · 2017
Cited alongside, same era.
Rotting bandits
N. Levine, K. Crammer, and S. Mannor · 2017
Cited alongside, same era.
Efficient contextual bandits in non-stationary worlds
H. Luo, C.-Y. Wei, A. Agarwal, and J. Langford · 2017
Cited alongside, same era.
Nearly optimal adaptive procedure for piecewise-stationary bandit: a change-point detection approach
Y. Cao, W. Zheng, B. Kveton, and Y. Xie · 2018
Later among the works it cites.
Learning to optimize under non-stationarity
W. C. Cheung, D. Simchi-Levi, and R. Zhu · 2018
Later among the works it cites.
Information directed sampling and bandits with heteroscedastic noise
J. Kirschner and A. Krause · 2018
Later among the works it cites.
Rotting bandits are no harder than stochastic ones
J. Seznec, A. Locatelli, A. Carpentier, A. Lazaric, and M. Valko · 2018
Later among the works it cites.
On abruptly-changing and slowly-varying multiarmed bandit problems
L. Wei and V. Srivatsva · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Non-stationary bandits with habituation and recovery dynamics
Y. Mintz, A. Aswani, P. Kaminsky, E. Flowers, and Y. Fukuoka · 2017
Cited alongside, same era.
Taming non-stationary bandits: A bayesian approach
V. Raj and S. Kalyani · 2017
Cited alongside, same era.
Adaptively tracking the best arm with an unknown number of distribution changes
P. Auer, P. Gajane, and R. Ortner · 2018
Cited alongside, same era.
Optimal exploration-exploitation in a multi-armed-bandit problem with non-stationary rewards
O. Besbes, Y. Gur, and A. Zeevi · 2018
Cited alongside, same era.
Later among the works it cites.
Learning contextual bandits in a non-stationary environment
Q. Wu, N. Iyer, and H. Wang · 2018
Later among the works it cites.
L. Besson and E. Kaufmann · 2019
Closest in time.
A new algorithm for non-stationary contextual bandits: Efficient, optimal, and parameter-free
Y. Chen, C.-W. Lee, H. Luo, and C.-Y. Wei · 2019
Closest in time.
Hedging the drift: Learning to optimize under non-stationarity
W. C. Cheung, D. Simchi-Levi, and R. Zhu · 2019
Closest in time.
Bandit Algorithms
T. Lattimore and C. Szepesvári · 2019
Closest in time.