Fetching the paper…
Reading the bibliography…
We study the non-stationary stochastic multiarmed bandit (MAB) problem and propose two generic algorithms, namely, the limited memory deterministic sequencing of exploration and exploitation (LM-DSEE) and the Sliding-Window Upper Confidence Bound# (SW-UCB#).
1963
Earlier work this paper cites.
J. R. Krebs, A. Kacelnik, and P. Taylor, “Test of optimal sampling by foraging great tits,” Nature , vol. 275, no. 5675, pp. 27–31, 1978
1978
Earlier work this paper cites.
T. L. Lai and H. Robbins, “Asymptotically efficient adaptive allocation rules,” Advances in Applied Mathematics , vol. 6, no. 1, pp. 4–22, 1985
1985
Earlier work this paper cites.
R. Agrawal, M. V. Hedge, and D. Teneketzis, “Asymptotically efficient adaptive allocation rules for the multi-armed bandit problem with switching cost,” IEEE Transactions on Automatic Control , vol. 33, no. 10, pp. 899–906, 1988
1988
Earlier work this paper cites.
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire, “The nonstochastic multiarmed bandit problem,” SIAM Journal on Computing , vol. 32, no. 1, pp. 48–77, 2002
2002
Earlier work this paper cites.
P. Auer, N. Cesa-Bianchi, and P. Fischer, “Finite-time analysis of the multiarmed bandit problem,” Machine Learning , vol. 47, no. 2, pp. 235–256, 2002
2002
Earlier work this paper cites.
2008
Earlier work this paper cites.
A. Anandkumar, N. Michael, A. K. Tang, and A. Swami, “Distributed algorithms for learning and cognitive medium access with logarithmic regret,” IEEE Journal on Selected Areas in Communications , vol. 29, no. 4, pp. 731–745, 2011
2011
Earlier work this paper cites.
N. Gupta, O.-C. Granmo, and A. Agrawala, “Thompson sampling for dynamic multi-armed bandits,” in International Conference on Machine Learning and Applications and Workshops , Honolulu, HI, Dec. 2011, pp. 484–489
2011
Cited alongside, same era.
S. Bubeck and N. Cesa-Bianchi, “Regret analysis of stochastic and nonstochastic multi-armed bandit problems,” Machine Learning , vol. 5, no. 1, pp. 1–122, 2012
2012
Cited alongside, same era.
V. Srivastava, P. Reverdy, and N. E. Leonard, “On optimal foraging and multi-armed bandits,” in Proceedings of the 51st Annual Allerton Conference on Communication, Control, and Computing , Monticello, IL, USA, 2013, pp. 494–499
2013
Cited alongside, same era.
M. Y. Cheung, J. Leighton, and F. S. Hover, “Autonomous mobile acoustic relay positioning as a multi-armed bandit with switching costs,” in IEEE/RSJ International Conference on Intelligent Robots and Systems , Tokyo, Japan, November 2013, pp. 3368–3373
2013
Cited alongside, same era.
2014
Later among the works it cites.
P. Reverdy, V. Srivastava, and N. E. Leonard, “Modeling human decision making in generalized Gaussian multiarmed bandits,” Proceedings of the IEEE , vol. 102, no. 4, pp. 544–571, 2014
2014
Later among the works it cites.
2015
Later among the works it cites.
N. Nayyar, D. Kalathil, and R. Jain, “On regret-optimal learning in decentralized multi-player multi-armed bandits,” IEEE Transactions on Control of Network Systems , vol. PP, no. 99, pp. 1–1, 2016
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Srivastava, F. Pasqualetti, and F. Bullo, “Stochastic surveillance strategies for spatial quickest detection,” The International Journal of Robotics Research , vol. 32, no. 12, pp. 1438–1458, 2013
2013
Cited alongside, same era.
S. Vakili, K. Liu, and Q. Zhao, “Deterministic sequencing of exploration and exploitation for multi-armed bandit problems,” IEEE Journal of Selected Topics in Signal Processing , vol. 7, no. 5, pp. 759–767, 2013
2013
Cited alongside, same era.
H. Liu, K. Liu, and Q. Zhao, “Learning in a changing world: Restless multiarmed bandit with unknown dynamics,” IEEE Transactions on Information Theory , vol. 59, no. 3, pp. 1902–1916, 2013
2013
Cited alongside, same era.
——, “Surveillance in an abruptly changing world via multiarmed bandits,” in IEEE Conference on Decision and Control , 2014, pp. 692–697
2014
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.