Fetching the paper…
Reading the bibliography…
We study an online decision making problem where on each round a learner chooses a list of items based on some side information, receives a scalar feedback value for each individual item, and a reward that is linearly related to this feedback.
The analysis of randomized and nonrandomized AIDS treatment trials using a new approach to causal inference in longitudinal studies
J. M. Robins · 1989
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John Lafferty, Andrew McCallum, and Fernando Pereira · 2001
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 2002
Earlier work this paper cites.
The on-line shortest path problem under partial monitoring
András György, Tamás Linder, Gábor Lugosi, and György Ottucsák · 2007
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2008
Earlier work this paper cites.
Search-based structured prediction
Hal Daumé III, John Langford, and Daniel Marcu · 2009
Earlier work this paper cites.
Algorithms for active learning
Daniel J Hsu · 2010
Earlier work this paper cites.
Non-stochastic bandit slate problems
Satyen Kale, Lev Reyzin, and Robert E Schapire · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert E Schapire · 2011
Cited alongside, same era.
Yahoo! learning to rank challenge overview
Olivier Chapelle and Yi Chang · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert E Schapire · 2011
Cited alongside, same era.
Efficient optimal learning for contextual bandits
Miroslav Dudík, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Cited alongside, same era.
User-Friendly Tail Bounds for Sums of Random Matrices
Joel A. Tropp · 2011
Cited alongside, same era.
Combinatorial bandits
Nicolo Cesa-Bianchi and Gábor Lugosi · 2012
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert E Schapire · 2014
Later among the works it cites.
Regret in online combinatorial optimization
Jean-Yves Audibert, Sébastien Bubeck, and Gábor Lugosi · 2014
Later among the works it cites.
Contextual combinatorial bandit and its application on diversified online recommendation
Lijing Qin, Shouyuan Chen, and Xiaoyan Zhu · 2014
Later among the works it cites.
Tight regret bounds for stochastic combinatorial semi-bandits
Branislav Kveton, Zheng Wen, Azin Ashkan, and Csaba Szepesvári · 2015
Closest in time.
Explore no more: Improved high-probability regret bounds for non-stochastic bandits
Gergely Neu · 2015
Closest in time.
Bistro: An efficient relaxation-based method for contextual bandits
Alexander Rakhlin and Karthik Sridharan · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Combinatorial multi-armed bandit: General framework and applications
Wei Chen, Yajun Wang, and Yang Yuan · 2013
Cited alongside, same era.
Hanson-wright inequality and sub-gaussian concentration
Mark Rudelson and Roman Vershynin · 2013
Cited alongside, same era.
Mslr: Microsoft learning to rank dataset
MSLR
Cited in the paper.
Closest in time.
Off-policy evaluation for slate recommendation
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miroslav Dudík, John Langford, Damien Jose, and Imed Zitouni · 2016
Closest in time.
Efficient algorithms for adversarial contextual learning
Vasilis Syrgkanis, Akshay Krishnamurthy, and Robert E Schapire · 2016
Closest in time.