Fetching the paper…
Reading the bibliography…
This paper studies online shortest path routing over multi-hop networks.
W. R. Thompson, “On the likelihood that one unknown probability exceeds another in view of the evidence of two samples,” Biometrika , vol. 25, no. 3/4, pp. 285–294, 1933
1933
Earlier work this paper cites.
H. Robbins, “Some aspects of the sequential design of experiments,” Bulletin of the American Mathematical Society , vol. 58, no. 5, pp. 527–535, 1952
1952
Earlier work this paper cites.
T. L. Lai and H. Robbins, “Asymptotically efficient adaptive allocation rules,” Advances in applied mathematics , vol. 6, no. 1, pp. 4–22, 1985
1985
Earlier work this paper cites.
V. Anantharam, P. Varaiya, and J. Walrand, “Asymptotically efficient allocation rules for the multiarmed bandit problem with multiple plays-Part I: IID rewards,” IEEE Transactions on Automatic Control , vol. 32, no. 11, pp. 968–976, 1987
1987
Earlier work this paper cites.
A. N. Burnetas and M. N. Katehakis, “Optimal adaptive policies for Markov decision processes,” Mathematics of Operations Research , vol. 22, no. 1, pp. 222–255, 1997
1997
Earlier work this paper cites.
T. L. Graves and T. L. Lai, “Asymptotically efficient adaptive choice of control laws in controlled Markov chains,” SIAM Journal on Control and Optimization , vol. 35, no. 3, pp. 715–743, 1997
1997
Earlier work this paper cites.
A. Sen and N. Balakrishnan, “Convolution of geometrics and a reliability problem,” Statistics & Probability Letters , vol. 43, no. 4, pp. 421–426, Jul. 1999
1999
Earlier work this paper cites.
P. Auer, N. Cesa-Bianchi, and P. Fischer, “Finite-time analysis of the multiarmed bandit problem,” Machine Learning , vol. 47, pp. 235–256, 2002
2002
Earlier work this paper cites.
B. Awerbuch and R. D. Kleinberg, “Adaptive routing with end-to-end feedback: Distributed learning and geometric approaches,” in Proceedings of the 36th Annual ACM Symposium on Theory of Computing (STOC) , 2004, pp. 45–53
2004
Earlier work this paper cites.
M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming . Wiley-Interscience, 2005
2005
Earlier work this paper cites.
A. György and G. Ottucsák, “Adaptive routing using expert advice,” The Computer Journal , vol. 49, no. 2, pp. 180–189, 2006
2006
Earlier work this paper cites.
A. György, T. Linder, G. Lugosi, and G. Ottucsák, “The on-line shortest path problem under partial monitoring,” Journal of Machine Learning Research , vol. 8, pp. 2369–2403, 2007
2007
Earlier work this paper cites.
2008
Earlier work this paper cites.
A. Shapiro, “Semi-infinite programming, duality, discretization and optimality conditions†,” Optimization , vol. 58, no. 2, pp. 133–161, 2009
2009
Cited alongside, same era.
Y. Gai, B. Krishnamachari, and R. Jain, “Learning multiuser channel allocations in cognitive radio networks: A combinatorial multi-armed bandit formulation,” in Proceedings of Symposium on New Frontiers in Dynamic Spectrum (DySPAN) , 2010
2010
Cited alongside, same era.
T. Jaksch, R. Ortner, and P. Auer, “Near-optimal regret bounds for reinforcement learning,” The Journal of Machine Learning Research , vol. 99, pp. 1563–1600, 2010
2010
Cited alongside, same era.
S. Filippi, O. Cappé, and A. Garivier, “Optimism in reinforcement learning and Kullback-Leibler divergence,” in Proceedings of the 48th Annual Allerton Conference on Communication, Control, and Computing , 2010, pp. 115–122
2010
Cited alongside, same era.
P. Joulani, A. György, and C. Szepesvári, “Online learning under delayed feedback,” in Proceedings of the 30th International Conference on Machine Learning (ICML) , 2013, pp. 1453–1461
2013
Closest in time.
Z. Zou, A. Proutiere, and M. Johansson, “Online shortest path routing: The value of information,” in Proceedings of American Control Conference (ACC) , Jun. 2014
2014
Closest in time.
A. Gopalan, S. Mannor, and Y. Mansour, “Thompson sampling for complex online problems,” in Proceedings of the 31st International Conference on Machine Learning (ICML) , 2014, pp. 100–108
2014
Closest in time.
J.-Y. Audibert, S. Bubeck, and G. Lugosi, “Regret in online combinatorial optimization,” Mathematics of Operations Research , vol. 39, no. 1, pp. 31–45, 2014
2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Garivier and O. Cappé, “The KL-UCB algorithm for bounded stochastic bandits and beyond,” in Proceedings of the 24th Conference On Learning Theory (COLT) , 2011
2011
Cited alongside, same era.
N. Cesa-Bianchi and G. Lugosi, “Combinatorial bandits,” Journal of Computer and System Sciences , vol. 78, no. 5, pp. 1404–1422, 2012
2012
Cited alongside, same era.
S. Bubeck, N. Cesa-Bianchi, and S. M. Kakade, “Towards minimax policies for online linear optimization with bandit feedback,” in Proceedings of the 25th Conference On Learning Theory (COLT) , 2012
2012
Cited alongside, same era.
Y. Gai, B. Krishnamachari, and R. Jain, “Combinatorial network optimization with unknown variables: Multi-armed bandits with linear rewards and individual observations,” IEEE/ACM Transactions on Networking , vol. 20, no. 5, pp. 1466–1478, 2012
2012
Cited alongside, same era.
K. Liu and Q. Zhao, “Adaptive shortest-path routing under unknown and stochastically varying link states,” in Proceedings of the 10th International Symposium on Modeling and Optimization in Mobile, Ad Hoc and Wireless Networks (WiOpt) , 2012, pp. 232–237
2012
Cited alongside, same era.
T. He, D. Goeckel, R. Raghavendra, and D. Towsley, “Endhost-based shortest path routing in dynamic networks,” in Proceedings of the 32nd IEEE International Conference on Computer Communications (INFOCOM) , 2013, pp. 2202–2210
2013
Cited alongside, same era.
W. Chen, Y. Wang, and Y. Yuan, “Combinatorial multi-armed bandit: General framework and applications,” in Proceedings of the 30th International Conference on Machine Learning (ICML) , 2013, pp. 151–159
2013
Cited alongside, same era.
G. Neu and G. Bartók, “An efficient algorithm for learning with semi-bandit feedback,” in Algorithmic Learning Theory (ALT) . Springer, 2013, pp. 234–248
2013
Cited alongside, same era.
2014
Closest in time.
2014
Closest in time.
S. Magureanu, R. Combes, and A. Proutiere, “Lipschitz bandits: Regret lower bounds and optimal algorithms,” in Proceedings of the 27th Conference on Learning Theory (COLT) , 2014
2014
Closest in time.
B. Kveton, Z. Wen, A. Ashkan, and C. Szepesvari, “Tight regret bounds for stochastic combinatorial semi-bandits,” in Proceedings of the 18th International Conference on Artificial Intelligence and Statistics (AISTATS) , 2015
2015
Closest in time.
R. Combes, M. S. Talebi, A. Proutiere, and M. Lelarge, “Combinatorial bandits revisited,” in Advances in Neural Information Processing Systems (NIPS) , 2015
2015
Closest in time.
Z. Wen, B. Kveton, and A. Ashkan, “Efficient learning in large-scale combinatorial semi-bandits,” in Proceedings of the 32nd International Conference on Machine Learning (ICML) , 2015, pp. 1113–1122
2015
Closest in time.
O. Brun, L. Wang, and E. Gelenbe, “Big data for autonomic intercontinental overlays,” IEEE Journal on Selected Areas in Communications , vol. 34, no. 3, pp. 575–583, 2016
2016
Closest in time.