Fetching the paper…
Reading the bibliography…
We consider the problem of distributed online learning with multiple players in multi-armed bandits (MAB) models.
D. Pollard, “Convergence of stochastic processes,” Springer , 1984
1984
Earlier work this paper cites.
T. Lai and H. Robbins, “Asymptotically efficient adaptive allocation rules,” Advances in Applied Mathematics , vol. 6, no. 1, pp. 4-22, 1985
1985
Earlier work this paper cites.
V. Anantharam, P. Varaiya, and J. Walrand, “Asymptotically efficient allocation rules for the multi-armed bandit problem with multiple plays - part i: i.i.d. rewards,” IEEE Transactions on Automatic Control , vol. 32, no. 11, pp. 968-975, November, 1987
1987
Earlier work this paper cites.
V. Anantharam, P. Varaiya, and J. Walrand, “Asymptotically efficient allocation rules for the multi-armed bandit problem with multiple plays - part ii: Markovian rewards,” IEEE Transactions on Automatic Control , vol. 32, no. 11, pp. 977-982, November 1987
1987
Earlier work this paper cites.
D. P. Bertsekas, “The auction algorithm: A distributed relaxation method for the assignment problem,” Annals of Operations Research , vol. 14, 1988
1988
Earlier work this paper cites.
D. P. Bertsekas, “Auction algorithms for network flow problems: A tutorial introduction,” Computational Optimization and Applications , vol. 1, pp. 7-66, 1992
1992
Earlier work this paper cites.
R. Agrawal, “Sample mean based index policies with ( O ( log n ) {O}(\log n) ) regret for the multi-armed bandit problem,” Advances in Applied Probability , Vol. 27, No. 4, pp. 1054-1078, 1995
1995
Earlier work this paper cites.
P. Lezaud, “Chernoff-type bound for finite markov chains,” Ann. Appl. Prob. , vol. 8, pp. 849-867, 1998
1998
Earlier work this paper cites.
C. Papadimitriou and J. Tsitsiklis, “The complexity of optimal queuing network control,” Mathematics of Operations Research , vol. 24, no. 2, pp. 293-305, May, 1999
1999
Cited alongside, same era.
P. Auer, N. Cesa-Bianchi, and P. Fischer, “Finite-time analysis of the multiarmed bandit problem,” Machine Learning , vol. 47, no. 2, pp. 235-256, 2002
2002
Cited alongside, same era.
E. Hossain and V. K. Bhargava, “Cognitive wireless communication networks,” Springer , 2007
2007
Cited alongside, same era.
M. Zavlanos, L. Spesivtsev, and G. J. Pappas, “A distributed auction algorithm for the assignment problem,” Proceedings of the IEEE Conference on Decision and Control , December, 2008
2008
Cited alongside, same era.
C. Tekin and M. Liu, “Online algorithms for the multi-armed bandit problem with markovian rewards,” Allerton Conference on Communication, Control, and Computing , October, 2010
W. Dai, Y. Gai, B. Krishnamachari, and Q. Zhao, “The non-bayesian restless multi-armed bandit: A case of near-logarithmic regret,” International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May, 2011
2011
Later among the works it cites.
Y. Gai, B. Krishnamachari, and M. Liu, “On the combinatorial multi-armed bandit problem with markovian rewards,” IEEE Global Communications Conference (GLOBECOM) , December, 2011
2011
Later among the works it cites.
A. Anandkumar, N. Michael, A. Tang, and A. Swami, “Distributed algorithms for learning and cognitive medium access with logarithmic regret,” IEEE JSAC on Advances in Cognitive Radio Networking and Communications , April, 2011
2011
Later among the works it cites.
Y. Gai and B. Krishnamachari, “Decentralized online learning algorithms for opportunistic spectrum access,” IEEE Global Communications Conference (GLOBECOM 2011) , December, 2011
2011
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2010
Cited alongside, same era.
K. Liu and Q. Zhao, “Distributed learning in multi-armed bandit with multiple players,” IEEE Transactions on Signal Processing , vol. 58, pp. 5667-5681, November, 2010
2010
Cited alongside, same era.
C. Tekin and M. Liu, “Online learning in opportunistic spectrum access: A restless bandit approach,” International Conference on Computer Communications (INFOCOM), Shanghai, China. , April 2011
2011
Cited alongside, same era.
H. Liu, K. Liu, and Q. Zhao, “Learning in a changing world: Restless multi-armed bandit with unknown dynamics,” IEEE Transactions on Information Theory , Submitted, November, 2011
2011
Cited alongside, same era.
K. Liu and Q. Zhao, “Multi-armed bandit problems with heavy tail reward distributions,” Allerton Conference on Communication, Control, and Computing , September, 2011
2011
Later among the works it cites.
W. Dai, Y. Gai, and B. Krishnamachari, “Efficient online learning for opportunistic spectrum access,” International Conference on Computer Communications (INFOCOM), Mini Conference, Orlando, USA , March, 2012
2012
Closest in time.
Y. Gai, B. Krishnamachari, and R. Jain, “Combinatorial network optimization with unknown variables: Multi-armed bandits with linear rewards and individual observations,” IEEE/ACM Trans. on Networking , to appear, 2012
2012
Closest in time.
C. Tekin and M. Liu, “Online learning of rested and restless bandits,,” IEEE Trans. on Information Theory , Submitted, 2012
2012
Closest in time.