Fetching the paper…
Reading the bibliography…
In a multi-armed bandit (MAB) problem a gambler needs to choose at each round of play one of K arms, each characterized by an unknown reward distribution.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. (1933) · 1933
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Robbins, H. (1952) · 1952
Earlier work this paper cites.
An analog of the minimax theorem for vector payoffs
Blackwell, D. (1956) · 1956
Earlier work this paper cites.
Approximation to bayes risk in repeated play, Contributions to the Theory of Games, Volume 3
Hannan, J. (1957) · 1957
Earlier work this paper cites.
Play the winner rule and the controlled clinical trials
Zelen, M. (1969) · 1969
Earlier work this paper cites.
A dynamic allocation index for the sequential design of experiments
Gittins, J. C. and D. M. Jones (1974) · 1974
Earlier work this paper cites.
Bandit processes and dynamic allocation indices (with discussion)
Gittins, J. C. (1979) · 1979
Earlier work this paper cites.
Arm acquiring bandits
Whittle, P. (1981) · 1981
Earlier work this paper cites.
Bandit problems: sequential allocation of experiments
Berry, D. A. and B. Fristedt (1985) · 1985
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L. and H. Robbins (1985) · 1985
Earlier work this paper cites.
Restless bandits: Activity allocation in a changing world
Whittle, P. (1988) · 1988
Earlier work this paper cites.
Multi-Armed Bandit Allocation Indices
Gittins, J. C. (1989) · 1989
Earlier work this paper cites.
The complexity of optimal queueing network control
Papadimitriou, C. H. and J. N. Tsitsiklis (1994) · 1994
Earlier work this paper cites.
Learning and strategic pricing
Bergemann, D. and J. Valimaki (1996) · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Freund, Y. and R. E. Schapire (1997) · 1997
Earlier work this paper cites.
Regret in the on-line decision problem
Foster, D. P. and R. V. Vohra (1999) · 1999
Cited alongside, same era.
Restless bandits, linear programming relaxations, and primal dual index heuristic
Bertsimas, D. and J. Nino-Mora (2000) · 2000
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Auer, P., N. Cesa-Bianchi, and P. Fischer (2002) · 2002
Cited alongside, same era.
The non-stochastic multi-armed bandit problem
Auer, P., N. Cesa-Bianchi, Y. Freund, and R. E. Schapire (2002) · 2002
Cited alongside, same era.
The value of knowing a demand curve: Bounds on regret for online posted-price auctions
Kleinberg, R. D. and T. Leighton (2003) · 2003
Cited alongside, same era.
Addaptive routing with end-to-end feedback: distributed learning and geometric approaches
Awerbuch, B. and R. D. Kleinberg (2004) · 2004
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S. and N. Cesa-Bianchi (2012) · 2012
Later among the works it cites.
Regret bounds for restless markov bandits
Ortner, R., D. Ryabko, P. Auer, and R. Munos (2012) · 2012
Later among the works it cites.
Online stochastic optimization under correlated bandit feedback
Azar, M. G., A. Lazaric, and E. Brunskill (2014) · 2014
Closest in time.
Stochastic multi-armed-bandit problem with non-stationary rewards
Besbes, O., Y. Gur, and A. Zeevi (2014) · 2014
Closest in time.
Contextual bandits with similarity information
Slivkins, A. (2014) · 2014
Closest in time.
Non-stationary stochastic optimization
Besbes, O., Y. Gur, and A. Zeevi (2015) · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The financing of innovation: Learning and stopping
Bergemann, D. and U. Hege (2005) · 2005
Cited alongside, same era.
Prediction, Learning, and Games
Cesa-Bianchi, N. and G. Lugosi (2006) · 2006
Cited alongside, same era.
Dynamic assortment with demand learning for seasonal consumer goods
Caro, F. and J. Gallien (2007) · 2007
Cited alongside, same era.
Approximation algorithms for partial-information based stochastic control with markovian rewards
Guha, S. and K. Munagala (2007) · 2007
Cited alongside, same era.
Bandits for taxonomies: A model-based approach
Pandey, S., D. Agarwal, D. Charkrabarti, and V. Josifovski (2007) · 2007
Cited alongside, same era.
Adapting to a changing environment: The brownian restless bandits
Slivkins, A. and E. Upfal (2008) · 2008
Cited alongside, same era.
Online optimization: Competing with dynamic comparators
Jadbabaie, A., A. Rakhlin, S. Shahrampour, and K. Sridharan (2015) · 2015
Closest in time.
Multi-armed bandits: Competing with optimal sequences
Karnin, Z. and O. Anava (2016) · 2016
Closest in time.
Tracking the best expert in non-stationary stochastic environments
Wei, C., Y. Hong, and C. Lu (2016) · 2016
Closest in time.
Rotting bandits
Levine, N., K. Crammer, and S. Mannor (2017) · 2017
Closest in time.
Nearly optimal adaptive procedure for piecewise-stationary bandit: a change-point detection approach
Cao, Y., W. Zheng, B. Kveton, and Y. Xie (2018) · 2018
Closest in time.
Learning to optimize under non-stationarity
Cheung, W. C., D. Simchi-Levi, and R. Zhu (2018) · 2018
Closest in time.
Efficient contextual bandits in nonstationary worlds
Luo, H., C. Wei, A. Agarwal, and J. Langford (2018) · 2018
Closest in time.
Dynamic regret of strongly adaptive methods
Zhang, L., T. Yang, and Z. Zhou (2018) · 2018
Closest in time.