Fetching the paper…
Reading the bibliography…
Online learning algorithms, widely used to power search and content optimization on the web, must balance exploration and exploitation, potentially sacrificing the experience of current users for information that will lead to better decisions in the future.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
An elementary proof of a theorem of Johnson and Lindenstrauss
Sanjoy Dasgupta and Anupam Gupta · 2003
Earlier work this paper cites.
Smoothed analysis of algorithms: Why the simplex algorithm usually takes polynomial time
Daniel A. Spielman and Shang-Hua Teng · 2004
Earlier work this paper cites.
The Epoch-Greedy Algorithm for Contextual Multi-armed Bandits
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
Upper and lower bounds for the normal distribution function, 2009
John D Cook · 2009
Earlier work this paper cites.
Introduction to Nonparametric Estimation
Alexandre B. Tsybakov · 2009
Earlier work this paper cites.
Regret bounds for sleeping experts and bandits
Robert Kleinberg, Alexandru Niculescu-Mizil, and Yogeshwer Sharma · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Earlier work this paper cites.
Nonparametric Bandits with Covariates
Philippe Rigollet and Assaf Zeevi · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Earlier work this paper cites.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Earlier work this paper cites.
The convex geometry of linear inverse problems
Venkat Chandrasekaran, Benjamin Recht, Pablo A Parrilo, and Alan S Willsky · 2012
Cited alongside, same era.
Fairness through awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel · 2012
Cited alongside, same era.
User-friendly tail bounds for sums of random matrices
Joel A Tropp · 2012
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Cited alongside, same era.
Incentivizing exploration
Peter Frazier, David Kempe, Jon M. Kleinberg, and Robert Kleinberg · 2014
Cited alongside, same era.
Implementing the “wisdom of the crowd”
Ilan Kremer, Yishay Mansour, and Motty Perry · 2014
Cited alongside, same era.
Making contextual decisions with low technical debt
Alekh Agarwal, Sarah Bird, Markus Cozowicz, Luong Hoang, John Langford, Stephen Lee, Jiaji Li, Dan Melamed, Gal Oshri, Oswaldo Ribas, Siddhartha Sen, and Alex Slivkins · 2017
Later among the works it cites.
Exploiting the natural exploration in contextual bandits
Hamsa Bastani, Mohsen Bayati, and Khashayar Khosravi · 2017
Later among the works it cites.
Crowdsourcing exploration
Kostas Bimpikis, Yiangos Papanastasiou, and Nicos Savva · 2017
Later among the works it cites.
L. Elisa Celis and Nisheeth K Vishnoi · 2017
Later among the works it cites.
Fair prediction with disparate impact: A study of bias in recidivism prediction instruments
Alexandra Chouldechova · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimal design for social learning
Yeon-Koo Che and Johannes Hörner · 2015
Cited alongside, same era.
Bayesian incentive-compatible bandit exploration
Yishay Mansour, Aleksandrs Slivkins, and Vasilis Syrgkanis · 2015
Cited alongside, same era.
Multiworld testing: A system for experimentation, learning, and decision-making
Alekh Agarwal, Sarah Bird, Markus Cozowicz, Miro Dudik, John Langford, Lihong Li, Luong Hoang, Dan Melamed, Siddhartha Sen, Robert Schapire, and Alex Slivkins · 2016
Cited alongside, same era.
Exploring or exploiting? Social and ethical implications of autonomous experimentation in AI
Sarah Bird, Solon Barocas, Kate Crawford, Fernando Diaz, and Hanna Wallach · 2016
Cited alongside, same era.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro · 2016
Cited alongside, same era.
Fairness in learning: Classic and contextual bandits
Matthew Joseph, Michael Kearns, Jamie Morgenstern, and Aaron Roth · 2016
Cited alongside, same era.
Meritocratic fairness for cross-population selection
Michael Kearns, Aaron Roth, and Zhiwei Steven Wu · 2017
Later among the works it cites.
Inherent trade-offs in the fair determination of risk scores
Jon Kleinberg, Sendhil Mullainathan, and Manish Raghavan · 2017
Later among the works it cites.
Calibrated fairness in bandits
Yang Liu, Goran Radanovic, Christos Dimitrakakis, Debmalya Mandal, and David C. Parkes · 2017
Later among the works it cites.
Practical evaluation and optimization of contextual bandit algorithms
Alberto Bietti, Alekh Agarwal, and John Langford · 2018
Closest in time.
Tail bounds for sums of geometric and exponential variables
Svante Janson · 2018
Closest in time.
A smoothed analysis of the greedy algorithm for the linear contextual bandit problem
Sampath Kannan, Jamie Morgenstern, Aaron Roth, Bo Waggoner, and Zhiwei Steven Wu · 2018
Closest in time.
Competing bandits: Learning under competition
Yishay Mansour, Aleksandrs Slivkins, and Zhiwei Steven Wu · 2018
Closest in time.