Fetching the paper…
Reading the bibliography…
Applications and systems are constantly faced with decisions that require picking from a set of actions based on contextual information.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 1995
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Y. Freund and R. E. Schapire · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
The Epoch-Greedy Algorithm for Contextual Multi-armed Bandits
J. Langford and T. Zhang · 2007
Earlier work this paper cites.
Controlled experiments on the web: survey and practical guide
R. Kohavi, R. Longbotham, D. Sommerfield, and R. M. Henne · 2009
Earlier work this paper cites.
The user and business impact of server delays, additional bytes, and http chunking in web search
E. Schurman and J. Brutlag · 2009
Earlier work this paper cites.
Agnostic active learning without constraints
A. Beygelzimer, J. Langford, Z. Tong, and D. J. Hsu · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Algorithms for Reinforcement Learning
C. Szepesvári · 2010
Earlier work this paper cites.
Efficient optimal leanring for contextual bandits
M. Dudik, D. Hsu, S. Kale, N. Karampatziakis, J. Langford, L. Reyzin, and T. Zhang · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and learning
M. Dudík, J. Langford, and L. Li · 2011
Cited alongside, same era.
Multi-Armed Bandit Allocation Indices
J. Gittins, K. Glazebrook, and R. Weber · 2011
Cited alongside, same era.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
L. Li, W. Chu, J. Langford, and X. Wang · 2011
Cited alongside, same era.
Contextual bandit learning with predictable rewards
A. Agarwal, M. Dudík, S. Kale, J. Langford, and R. E. Schapire · 2012
Cited alongside, same era.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
Sample-efficient nonstationary policy evaluation for contextual bandits
M. Dudík, D. Erhan, J. Langford, and L. Li · 2012
Cited alongside, same era.
Scaling distributed machine learning with the parameter server
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su · 2014
Later among the works it cites.
Machine learning: The high-interest credit card of technical debt
D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner, V. Chaudhary, and M. Young · 2014
Later among the works it cites.
Ice: enabling non-experts to build models interactively for large-scale lopsided problems
P. Simard, D. Chickering, A. Lakshmiratan, D. Charles, L. Bottou, C. G. J. Suarez, D. Grangier, S. Amershi, J. Verwey, and J. Suh · 2014
Later among the works it cites.
Minerva: A scalable and highly efficient training platform for deep learning
M. Wang, T. Xiao, J. Li, J. Zhang, C. Hong, and Z. Zhang · 2014
Later among the works it cites.
Personalizing linkedin feed
D. Agarwal, B.-C. Chen, Q. He, Z. Hua, G. Lebanon, Y. Ma, P. Shivaswamy, H.-P. Tseng, J. Yang, and L. Zhang · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Field Experiments: Design, Analysis, and Interpretation
A. S. Gerber and D. P. Green · 2012
Cited alongside, same era.
Counterfactual reasoning and learning systems: The example of computational advertising
L. Bottou, J. Peters, J. Quinonero-Candela, D. X. Charles, D. M. Chickering, E. Portugaly, D. Ray, P. Simard, and E. Snelson · 2013
Cited alongside, same era.
A reliable effective terascale linear learning system
A. Agarwal, O. Chapelle, M. Dudík, and J. Langford · 2014
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. Schapire · 2014
Cited alongside, same era.
Theory of disagreement-based active learning
S. Hanneke · 2014
Cited alongside, same era.
https://aws.amazon.com/machine-learning/
Amazon Machine Learning - Predictive Analytics with AWS
Cited in the paper.
NEXT: A system for real-world development, evaluation, and application of active learning
K. G. Jamieson, L. Jain, C. Fernandez, N. J. Glattard, and R. Nowak · 2015
Later among the works it cites.
Online controlled experiments and a/b tests
R. Kohavi and R. Longbotham · 2015
Later among the works it cites.
Efficient contextual semi-bandit learning
A. Krishnamurthy, A. Agarwal, and M. Dudík · 2015
Later among the works it cites.
White paper,
Multi-world testing: A system for experimentation, learning, and decision-making · 2016
Closest in time.
Clipper: A low-latency online prediction serving system
D. Crankshaw, X. Wang, G. Zhou, M. J. Franklin, J. E. Gonzalez, and I. Stoica · 2017
Closest in time.