Fetching the paper…
Reading the bibliography…
Multi-armed bandit problems are considered as a paradigm of the trade-off between exploring the environment to find profitable actions and exploiting what is already known.
Asymptotically efficient adaptive allocation rules
T. L. Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Restless bandits: activity allocation in a changing world
P. Whittle · 1988
Earlier work this paper cites.
Sample mean based index policies with O ( log n ) O(\log n) regret for the multi-armed bandit problem
R. Agrawal · 1995
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E. Schapire · 1995
Earlier work this paper cites.
A probabilistic theory of pattern recognition , volume 31 of Applications of Mathematics (New York)
L. Devroye, L. Györfi, and G. Lugosi · 1996
Earlier work this paper cites.
Tracking the best expert
M. Herbster and M.K. Warmuth · 1998
Earlier work this paper cites.
On prediction of individual sequences
Nicolò Cesa-Bianchi and Gábor Lugosi · 1999
Earlier work this paper cites.
Finite-time lower bounds for the two-armed bandit problem
S. R. Kulkarni and G. Lugosi · 2000
Cited alongside, same era.
The nonstochastic multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R.E. Schapire · 2002
Cited alongside, same era.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer · 2002
Cited alongside, same era.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gabor Lugosi · 2006
Cited alongside, same era.
Regret minimization under partial monitoring
Nicolò Cesa-Bianchi, Gábor Lugosi, and Gilles Stoltz · 2006
Cited alongside, same era.
Discounted UCB
L. Kocsis and C. Szepesvári · 2006
Later among the works it cites.
Tuning bandit algorithms in stochastic environments
J-Y. Audibert, R. Munos, and A. Szepesvári · 2007
Later among the works it cites.
Cognitive medium access: Exploration, exploitation and competition, 2007
L. Lai, H. El Gamal, H. Jiang, and H. V. Poor · 2007
Later among the works it cites.
Competing with typical compound actions, 2008
Nicolò Cesa-Bianchi, Gábor Lugosi, and Gilles Stoltz · 2008
Closest in time.
Reinforcement learning and evolutionary algorithms for non-stationary multi-armed bandit problems
D. E. Koulouriotis and A. Xanthopoulos · 2008
Closest in time.
Adapting to a changing environment: the brownian restless bandits, 2008
A. Slivkins and E. Upfal · 2008
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-armed bandit, dynamic environments and meta-bandits, 2006
C. Hartland, S. Gelly, N. Baskiotis, O. Teytaud, and M. Sebag · 2006
Cited alongside, same era.