Fetching the paper…
Reading the bibliography…
Many sequential decision-making problems in communication networks can be modeled as contextual bandit problems, which are natural extensions of the well-known multi-armed bandit problem.
W. Hoeffding, “Probability inequalities for sums of bounded random variables,” Journal of the American statistical association
1963
Earlier work this paper cites.
T. L. Lai and H. Robbins, “Asymptotically efficient adaptive allocation rules,” Advances in applied mathematics
1985
Earlier work this paper cites.
R. Agrawal, “Sample mean based index policies with o (log n) regret for the multi-armed bandit problem,” Advances in Applied Probability
1995
Earlier work this paper cites.
A. N. Burnetas and M. N. Katehakis, “Optimal adaptive policies for markov decision processes,” Mathematics of Operations Research
1997
Earlier work this paper cites.
P. Auer, N. Cesa-Bianchi, and P. Fischer, “Finite-time analysis of the multiarmed bandit problem,” Machine learning
2002
Earlier work this paper cites.
P. Auer, “Using confidence bounds for exploitation-exploration trade-offs,” The Journal of Machine Learning Research
2003
Earlier work this paper cites.
J. Langford and T. Zhang, “The epoch-greedy algorithm for multi-armed bandits with side information,” in Advances in neural information processing systems
2008
Earlier work this paper cites.
T. Lu, D. Pál, and M. Pál, “Contextual multi-armed bandits,” in International Conference on Artificial Intelligence and Statistics
2010
Cited alongside, same era.
L. Li, W. Chu, J. Langford, and R. E. Schapire, “A contextual-bandit approach to personalized news article recommendation,” in Proceedings of the 19th international conference on World wide web
2010
Cited alongside, same era.
W. Chu, L. Li, L. Reyzin, and R. E. Schapire, “Contextual bandits with linear payoff functions,” in International Conference on Artificial Intelligence and Statistics
2011
Cited alongside, same era.
M. Dudik, D. Hsu, S. Kale, N. Karampatziakis, J. Langford, L. Reyzin, and T. Zhang, “Efficient optimal learning for contextual bandits,” in Conference on Uncertainty in Artificial Intelligence
2011
Cited alongside, same era.
S. Agrawal and N. Goyal, “Thompson sampling for contextual bandits with linear payoffs,” in International Conference on Machine Learning
2013
Later among the works it cites.
Y. Gai and B. Krishnamachari, “Distributed stochastic online learning policies for opportunistic spectrum access,” Signal Processing, IEEE Transactions on
2014
Later among the works it cites.
A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. E. Schapire, “Taming the monster: A fast and simple algorithm for contextual bandits,” in International Conference on Machine Learning
2014
Later among the works it cites.
A. Slivkins, “Contextual bandits with similarity information,” The Journal of Machine Learning Research
2014
Later among the works it cites.
A. Badanidiyuru, J. Langford, and A. Slivkins, “Resourceful contextual bandits,” in Conference on Learning Theory
2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Gai, B. Krishnamachari, and R. Jain, “Combinatorial network optimization with unknown variables: Multi-armed bandits with linear rewards and individual observations,” IEEE/ACM Transactions on Networking (TON)
2012
Cited alongside, same era.
S. Vakili, K. Liu, and Q. Zhao, “Deterministic sequencing of exploration and exploitation for multi-armed bandit problems,” Selected Topics in Signal Processing, IEEE Journal of
2013
Cited alongside, same era.
Later among the works it cites.
H. Wu, R. Srikant, X. Liu, and C. Jiang, “Algorithms with logarithmic or sublinear regret for constrained contextual bandits,” in Advances in Neural Information Processing Systems
2015
Later among the works it cites.