Fetching the paper…
Reading the bibliography…
Multi-armed bandit (MAB) is a class of online learning problems where a learning agent aims to maximize its expected cumulative reward while repeatedly selecting to pull arms with unknown reward distributions.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. (1933) · 1933
Earlier work this paper cites.
Continuous inspection schemes
Page, E. S. (1954) · 1954
Earlier work this paper cites.
A generalized likelihood ratio approach to the detection and estimation of jumps in linear systems
Willsky, A. and Jones, H. (1976) · 1976
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L. and Robbins, H. (1985) · 1985
Earlier work this paper cites.
Sequential analysis: tests and confidence intervals
Siegmund, D. (1985) · 1985
Earlier work this paper cites.
Detection of abrupt changes: theory and application
Basseville, M., Nikiforov, I. V., et al. (1993) · 1993
Earlier work this paper cites.
The weighted majority algorithm
Littlestone, N. and Warmuth, M. K. (1994) · 1994
Earlier work this paper cites.
Tracking the best expert
Herbster, M. and Warmuth, M. K. (1998) · 1998
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P. (2002) · 2002
Earlier work this paper cites.
Multi-armed bandit algorithms and empirical evaluation
Vermorel, J. and Mohri, M. (2005) · 2005
Earlier work this paper cites.
Prediction, learning, and games
Cesa-Bianchi, N. and Lugosi, G. (2006) · 2006
Cited alongside, same era.
Discounted ucb
Kocsis, L. and Szepesvári, C. (2006) · 2006
Cited alongside, same era.
Change point detection and meta-bandits for online learning in dynamic environments
Hartland, C., Baskiotis, N., Gelly, S., Sebag, M., and Teytaud, O. (2007) · 2007
Cited alongside, same era.
Dynamic spectrum access with non-stationary multi-armed bandit
Alaya-Feki, A. B. H., Moulines, E., and LeCornec, A. (2008) · 2008
Cited alongside, same era.
On upper-confidence bound policies for non-stationary bandit problems
Garivier, A. and Moulines, E. (2008) · 2008
Cited alongside, same era.
A case study of behavior-driven conjoint analysis on yahoo!: front page today module
Chu, W., Park, S.-T., Beaupre, T., Motgi, N., Phadke, A., Chakraborty, S., and Zachariah, J. (2009) · 2009
A contextual-bandit algorithm for mobile context-aware recommender system
Bouneffouf, D., Bouzeghoub, A., and Gançarski, A. L. (2012) · 2012
Later among the works it cites.
Managing advertising campaigns?an approximate planning approach
Girgin, S., Mary, J., Preux, P., and Nicol, O. (2012) · 2012
Later among the works it cites.
Thompson sampling in switching environments with bayesian online change detection
Mellor, J. and Shapiro, J. (2013) · 2013
Later among the works it cites.
Stochastic multi-armed-bandit problem with non-stationary rewards
Besbes, O., Gur, Y., and Zeevi, A. (2014) · 2014
Later among the works it cites.
Matroid bandits: Fast combinatorial optimization with learning
Kveton, B., Wen, Z., Ashkan, A., Eydgahi, H., and Eriksson, B. (2014) · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Piecewise-stationary bandit problems with side observations
Yu, J. Y. and Mannor, S. (2009) · 2009
Cited alongside, same era.
Sequential change-point detection when the pre-and post-change parameters are unknown
Lai, T. L. and Xing, H. (2010) · 2010
Cited alongside, same era.
On upper-confidence bound policies for switching bandit problems
Garivier, A. and Moulines, E. (2011) · 2011
Cited alongside, same era.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Li, L., Chu, W., Langford, J., and Wang, X. (2011) · 2011
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P. (2002a)
Cited in the paper.
The nonstochastic multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E. (2002b)
Cited in the paper.
Allesiardo, R. and Féraud, R. (2015) · 2015
Later among the works it cites.
Multi-armed bandit models for the optimal design of clinical trials: benefits and challenges
Villar, S. S., Bowden, J., and Wason, J. (2015) · 2015
Later among the works it cites.
A change-detection based framework for piecewise-stationary multi-armed bandit problem
Liu, F., Lee, J., and Shroff, N. (2017) · 2017
Later among the works it cites.
Customer acquisition via display advertising using multi-armed bandit experiments
Schwartz, E. M., Bradlow, E. T., and Fader, P. S. (2017) · 2017
Later among the works it cites.