Fetching the paper…
Reading the bibliography…
In the Multi-Armed Bandit (MAB) problem, there is a given set of arms with unknown reward models.
H. Robbins, “Some Aspects of the Sequential Design of Experiments,”
1952
Earlier work this paper cites.
W. Hoeffding, “Probability Inequalities for Sums of Bounded Random Variables,”
1963
Earlier work this paper cites.
T. Santner, A. Tamhane,
1984
Earlier work this paper cites.
T. Lai and H. Robbins, “Asymptotically Efficient Adaptive Allocation Rules,”
1985
Earlier work this paper cites.
V. Anantharam, P. Varaiya, J. Walrand, “Asymptotically Efficient Allocation Rules for the Multiarmed Bandit Problem with Multiple Plays-Part II: Markovian Rewards,”
1987
Earlier work this paper cites.
R. Agrawal, “Sample Mean Based Index Policies with
1995
Earlier work this paper cites.
R. Agrawal, “The Continuum-Armed Bandit Problem,”
1995
Earlier work this paper cites.
Y. Ren, H. Liang “On the best constant in Marcinkiewicz-Zygmund inequality,”
2001
Cited alongside, same era.
P. Auer, N. Cesa-Bianchi, P. Fischer, “Finite-time Analysis of the Multiarmed Bandit Problem,”
2002
Cited alongside, same era.
S. Bubeck, N. Cesa-Bianchi, G. Lugosi, “Bandits with heavy tail,”
2002
Cited alongside, same era.
P. Chareka, O. Chareka, S. Kennendy, “Locally Sub-Gaussian Random Variable and the Stong Law of Large Numbers,”
2006
Cited alongside, same era.
A. Mahajan and D. Teneketzis, “Multi-armed Bandit Problems,”
2007
Cited alongside, same era.
K. Liu, Q. Zhao, “Distributed Learning in Multi-Armed Bandit with Multiple Players,”
A. Anandkumar, N. Michael, A.K. Tang, A. Swami, “Distributed Algorithms for Learning and Cognitive Medium Access with Logarithmic Regret,”
2011
Closest in time.
Y. Gai and B. Krishnamachari, “Decentralized Online Learning Algorithms for Opportunistic Spectrum Access,”
2011
Closest in time.
C. Tekin, M. Liu, “Performance and Convergence of Multiuser Online Learning,”
2011
Closest in time.
Y. Gai and B. Krishnamachari, “Decentralized Online Learning Algorithms for Opportunistic Spectrum Access,” Technical Report, March, 2011. Available at http://anrg.usc.edu/www/publications/papers/DMAB2011.pdf
2011
Closest in time.
K. Liu and Q. Zhao, “Adaptive Shortest-Path Routing under Unknown and Stochastically Varying Link States,”
2012
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2010
Cited alongside, same era.
C. Tekin, M. Liu, “Online Algorithms for the Multi-Armed Bandit Problem With Markovian Rewards,”
2010
Cited alongside, same era.
H. Liu, K. Liu, and Q. Zhao, “Learning in A Changing World: Restless Multi-Armed Bandit with Unknown Dynamics,” to appear in
Cited in the paper.
D. Kalathil, N. Nayyar, R. Jain, “Decentralized learning for multi-player multi-armed bandits,”
2012
Closest in time.