Fetching the paper…
Reading the bibliography…
The problem of multi-armed bandits (MAB) asks to make sequential decisions while balancing between exploitation and exploration, and have been successfully applied to a wide range of practical scenarios.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Herbert Robbins · 1985
Earlier work this paper cites.
Asymptotically efficient allocation rules for the multiarmed bandit problem with multiple plays-part i: Iid rewards
Venkatachalam Anantharam, Pravin Varaiya, and Jean Walrand · 1987
Earlier work this paper cites.
Sample mean based index policies by o (log n) regret for the multi-armed bandit problem
Rajeev Agrawal · 1995
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2008
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Sarah Filippi, Olivier Cappe, Aurélien Garivier, and Csaba Szepesvári · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert E Schapire · 2011
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Olivier Chapelle and Lihong Li · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert E Schapire · 2011
Cited alongside, same era.
Contextual gaussian process bandit optimization
Andreas Krause and Cheng S Ong · 2011
Cited alongside, same era.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Cited alongside, same era.
Linear submodular bandits and their application to diversified retrieval
Yisong Yue and Carlos Guestrin · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Cited alongside, same era.
Networked bandits with disjoint linear payoffs
Meng Fang and Dacheng Tao · 2014
Later among the works it cites.
Combinatorial partial monitoring game with linear feedback and its applications
Tian Lin, Bruno D Abrahao, Robert D Kleinberg, John Lui, and Wei Chen · 2014
Later among the works it cites.
Contextual combinatorial bandit and its application on diversified online recommendation
Lijing Qin, Shouyuan Chen, and Xiaoyan Zhu · 2014
Later among the works it cites.
Tight regret bounds for stochastic combinatorial semi-bandits
Branislav Kveton, Zheng Wen, Azin Ashkan, and Csaba Szepesvari · 2015
Later among the works it cites.
Combinatorial multi-armed bandit and its extension to probabilistically triggered arms
Wei Chen, Yajun Wang, Yang Yuan, and Qinshi Wang · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Combinatorial network optimization with unknown variables: Multi-armed bandits with linear rewards and individual observations
Yi Gai, Bhaskar Krishnamachari, and Rahul Jain · 2012
Cited alongside, same era.
Combinatorial multi-armed bandit: General framework, results and applications
Wei Chen, Yajun Wang, and Yang Yuan · 2013
Cited alongside, same era.
Automatic ad format selection via contextual bandits
Liang Tang, Romer Rosales, Ajit Singh, and Deepak Agarwal · 2013
Cited alongside, same era.
Combinatorial pure exploration of multi-armed bandits
Shouyuan Chen, Tian Lin, Irwin King, Michael R Lyu, and Wei Chen · 2014
Cited alongside, same era.
Abbas Kazerouni, Mohammad Ghavamzadeh, and Benjamin Van Roy · 2016
Later among the works it cites.
Conservative bandits
Yifan Wu, Roshan Shariff, Tor Lattimore, and Csaba Szepesvári · 2016
Later among the works it cites.
Linear submodular bandits with a knapsack constraint
Baosheng Yu, Meng Fang, and Dacheng Tao · 2016
Later among the works it cites.
Diffusion independent semi-bandit influence maximization
Sharan Vaswani, Branislav Kveton, Zheng Wen, Mohammad Ghavamzadeh, Laks Lakshmanan, and Mark Schmidt · 2017
Later among the works it cites.
Efficient ordered combinatorial semi-bandits for whole-page recommendation
Yingfei Wang, Hua Ouyang, Chu Wang, Jianhui Chen, Tsvetan Asamov, and Yi Chang · 2017
Later among the works it cites.