Fetching the paper…
Reading the bibliography…
This paper studies the evaluation of policies that recommend an ordered set of items (e.g., a ranking) based on some context---a common scenario in web search, ads, and recommendation.
A generalization of sampling without replacement from a finite universe
Daniel G Horvitz and Donovan J Thompson · 1952
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Learning to rank using gradient descent
Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender · 2005
Earlier work this paper cites.
The price of bandit information for online optimization
Varsha Dani, Thomas P. Hayes, and Sham M. Kakade · 2008
Earlier work this paper cites.
A user browsing model to predict search engine click data from past observations
Georges E. Dupret and Benjamin Piwowarski · 2008
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2008
Earlier work this paper cites.
Exploration scavenging
John Langford, Alexander Strehl, and Jennifer Wortman · 2008
Earlier work this paper cites.
The matrix cookbook
Kaare Brandt Petersen, Michael Syskind Pedersen, et al · 2008
Earlier work this paper cites.
A dynamic Bayesian network click model for web search ranking
Olivier Chapelle and Ya Zhang · 2009
Earlier work this paper cites.
Expected reciprocal rank for graded relevance
Olivier Chapelle, Donald Metlzer, Ya Zhang, and Pierre Grinspan · 2009
Earlier work this paper cites.
Click chain model in web search
Fan Guo, Chao Liu, Anitha Kannan, Tom Minka, Michael Taylor, Yi-Min Wang, and Christos Faloutsos · 2009
Earlier work this paper cites.
Controlled experiments on the web: survey and practical guide
Ron Kohavi, Roger Longbotham, Dan Sommerfield, and Randal M Henne · 2009
Cited alongside, same era.
Parametric bandits: The generalized linear case
Sarah Filippi, Olivier Cappe, Aurélien Garivier, and Csaba Szepesvári · 2010
Cited alongside, same era.
Non-stochastic bandit slate problems
Satyen Kale, Lev Reyzin, and Robert E Schapire · 2010
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Cited alongside, same era.
Linearly parameterized bandits
Paat Rusmevichientong and John N Tsitsiklis · 2010
Cited alongside, same era.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert E Schapire · 2011
Cited alongside, same era.
Introducing LETOR 4.0 datasets
Tao Qin and Tie-Yan Liu · 2013
Later among the works it cites.
Doubly robust policy evaluation and optimization
Miroslav Dudík, Dumitru Erhan, John Langford, and Lihong Li · 2014
Later among the works it cites.
Contextual combinatorial bandit and its application on diversified online recommendation
Lijing Qin, Shouyuan Chen, and Xiaoyan Zhu · 2014
Later among the works it cites.
Tight regret bounds for stochastic combinatorial semi-bandits
Branislav Kveton, Zheng Wen, Azin Ashkan, and Csaba Szepesvári · 2015
Later among the works it cites.
Toward predicting the outcome of an a/b experiment for search relevance
Lihong Li, Imed Zitouni, and Jin Young Kim · 2015
Later among the works it cites.
Counterfactual risk minimization: Learning from logged bandit feedback
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Cited alongside, same era.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolò Cesa-Bianchi · 2012
Cited alongside, same era.
Combinatorial bandits
Nicolo Cesa-Bianchi and Gábor Lugosi · 2012
Cited alongside, same era.
Training efficient tree-based models for document ranking
Nima Asadi and Jimmy Lin · 2013
Cited alongside, same era.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis Charles, Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Cited alongside, same era.
Adith Swaminathan and Thorsten Joachims · 2015
Later among the works it cites.
A cross-benchmark comparison of 87 learning to rank methods
Niek Tax, Sander Bockting, and Djoerd Hiemstra · 2015
Later among the works it cites.
Online evaluation for information retrieval
Katja Hofmann, Lihong Li, Filip Radlinski, et al · 2016
Closest in time.
Efficient contextual semi-bandit learning
Akshay Krishnamurthy, Alekh Agarwal, and Miroslav Dudík · 2016
Closest in time.
Beyond ranking: Optimizing whole-page presentation
Yue Wang, Dawei Yin, Luo Jie, Pengyuan Wang, Makoto Yamada, Yi Chang, and Qiaozhu Mei · 2016
Closest in time.
Optimal and adaptive off-policy evaluation in contextual bandits
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudik · 2017
Closest in time.