Fetching the paper…
Reading the bibliography…
Cascading bandit (CB) is a popular model for web search and online advertising, where an agent aims to learn the $K$ most attractive items out of a ground set of size $L$ during the interaction with a user.
Besson, L. and Kaufmann, E. (2019) · 1902
Earlier work this paper cites.
Procedures for reacting to a change in distribution
Lorden, G. et al. (1971) · 1908
Earlier work this paper cites.
Stochastic bandits with delay-dependent payoffs
Cella, L. and Cesa-Bianchi, N. (2019) · 1910
Earlier work this paper cites.
Continuous inspection schemes
Page, E. S. (1954) · 1954
Earlier work this paper cites.
Inference about the change-point from cumulative sum tests
Hinkley, D. V. (1971) · 1971
Earlier work this paper cites.
A generalized likelihood ratio approach to the detection and estimation of jumps in linear systems
Willsky, A. and Jones, H. (1976) · 1976
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L. and Robbins, H. (1985) · 1985
Earlier work this paper cites.
Optimal stopping times for detecting changes in distributions
Moustakides, G. V. et al. (1986) · 1986
Earlier work this paper cites.
Restless bandits: Activity allocation in a changing world
Whittle, P. (1988) · 1988
Earlier work this paper cites.
The weighted majority algorithm
Littlestone, N. and Warmuth, M. K. (1994) · 1994
Earlier work this paper cites.
Using the generalized likelihood ratio statistic for sequential detection of a change-point
Siegmund, D. and Venkatraman, E. (1995) · 1995
Earlier work this paper cites.
Multihypothesis sequential probability ratio tests. i. asymptotic optimality
Draglia, V., Tartakovsky, A. G., and Veeravalli, V. V. (1999) · 1999
Earlier work this paper cites.
Combinatorial semi-bandit in the non-stationary environment
Chen, W., Wang, L., Zhao, H., and Zheng, K. (2020) · 2002
Earlier work this paper cites.
Optimal and asymptotically optimal cusum rules for change point detection in the brownian motion model with multiple alternatives
Hadjiliadis, O. and Moustakides, V. (2006) · 2006
Earlier work this paper cites.
Discounted ucb
Kocsis, L. and Szepesvári, C. (2006) · 2006
Earlier work this paper cites.
Change point detection and meta-bandits for online learning in dynamic environments
Hartland, C., Baskiotis, N., Gelly, S., Sebag, M., and Teytaud, O. (2007) · 2007
Earlier work this paper cites.
An experimental comparison of click position-bias models
Craswell, N., Zoeter, O., Taylor, M., and Ramsey, B. (2008) · 2008
Cited alongside, same era.
A user browsing model to predict search engine click data from past observations
Dupret, G. E. and Piwowarski, B. (2008) · 2008
Cited alongside, same era.
Piecewise-stationary bandit problems with side observations
Yu, J. Y. and Mannor, S. (2009) · 2009
Cited alongside, same era.
Learning more powerful test statistics for click-based retrieval evaluation
Yue, Y., Gao, Y., Chapelle, O., Zhang, Y., and Joachims, T. (2010) · 2010
Cited alongside, same era.
On upper-confidence bound policies for switching bandit problems
Garivier, A. and Moulines, E. (2011) · 2011
Cited alongside, same era.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Li, L., Chu, W., Langford, J., and Wang, X. (2011) · 2011
Improving regret bounds for combinatorial semi-bandits with probabilistically triggered arms and its applications
Wang, Q. and Chen, W. (2017) · 2017
Later among the works it cites.
Online learning to rank in stochastic click models
Zoghi, M., Tunys, T., Ghavamzadeh, M., Kveton, B., Szepesvari, C., and Wen, Z. (2017) · 2017
Later among the works it cites.
Mixture martingales revisited with applications to sequential tests and confidence intervals
Kaufmann, E. and Koolen, W. (2018) · 2018
Later among the works it cites.
A change-detection based framework for piecewise-stationary multi-armed bandit problem
Liu, F., Lee, J., and Shroff, N. (2018) · 2018
Later among the works it cites.
Stochastic bandits robust to adversarial corruptions
Lykouris, T., Mirrokni, V., and Paes Leme, R. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Combinatorial bandits
Cesa-Bianchi, N. and Lugosi, G. (2012) · 2012
Cited alongside, same era.
Kullback–leibler upper confidence bounds for optimal sequential allocation
Cappé, O., Garivier, A., Maillard, O.-A., Munos, R., Stoltz, G., et al. (2013) · 2013
Cited alongside, same era.
Sequential analysis: tests and confidence intervals
Siegmund, D. (2013) · 2013
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Besbes, O., Gur, Y., and Zeevi, A. (2014) · 2014
Cited alongside, same era.
Combinatorial bandits revisited
Combes, R., Shahi, M. S. T. M., Proutiere, A., et al. (2015) · 2015
Cited alongside, same era.
Cascading bandits: Learning to rank in the cascade model
Kveton, B., Szepesvari, C., Wen, Z., and Ashkan, A. (2015) · 2015
Cited alongside, same era.
On analyzing user preference dynamics with temporal social networks
Pereira, F. S., Gama, J., de Amo, S., and Oliveira, G. M. (2018) · 2018
Later among the works it cites.
More adaptive algorithms for adversarial bandits
Wei, C.-Y. and Luo, H. (2018) · 2018
Later among the works it cites.
On abruptly-changing and slowly-varying multiarmed bandit problems
Wei, L. and Srivatsva, V. (2018) · 2018
Later among the works it cites.
Adaptively tracking the best bandit arm with an unknown number of distribution changes
Auer, P., Gajane, P., and Ortner, R. (2019) · 2019
Closest in time.
Nearly optimal adaptive procedure with change detection for piecewise-stationary bandit
Cao, Y., Wen, Z., Kveton, B., and Xie, Y. (2019) · 2019
Closest in time.
A thompson sampling algorithm for cascading bandits
Cheung, W. C., Tan, V., and Zhong, Z. (2019) · 2019
Closest in time.
When people change their mind: Off-policy evaluation in non-stationary recommendation environments
Jagerman, R., Markov, I., and de Rijke, M. (2019) · 2019
Closest in time.
Bandit online learning with unknown delays
Li, B., Chen, T., and Giannakis, G. B. (2019) · 2019
Closest in time.
Cascading non-stationary bandits: Online learning to rank in the non-stationary cascade model
Li, C. and de Rijke, M. (2019) · 2019
Closest in time.
A near-optimal change-detection based algorithm for piecewise-stationary combinatorial semi-bandits
Zhou, H., Wang, L., Varshney, L. R., and Lim, E.-P. (2020) · 2020
Closest in time.