Fetching the paper…
Reading the bibliography…
In a multi-armed bandit problem, an online algorithm chooses from a set of strategies in a sequence of trials so as to maximize the total payoff of the chosen strategies.
Contribution à la topologie des ensembles dénombrables
S. Mazurkiewicz and W. Sierpinski · 1920
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
A comparison of signalling alphabets
E. N. Gilbert · 1952
Earlier work this paper cites.
Some Aspects of the Sequential Design of Experiments
Herbert Robbins · 1952
Earlier work this paper cites.
Estimate of the number of signals in error correcting codes
R. R. Varshamov · 1957
Earlier work this paper cites.
Über unendliche, lineare Punktmannichfaltigkeiten, 4
G. Cantor · 1966
Earlier work this paper cites.
Bandit problems: sequential allocation of experiments
Donald Berry and Bert Fristedt · 1985
Earlier work this paper cites.
Asymptotically efficient Adaptive Allocation Rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Elements of Information Theory
Thomas M. Cover and Joy A. Thomas · 1991
Earlier work this paper cites.
Fractal, Chaos and Power Laws: Minutes from an Infinite Paradise
Manfred Schroeder · 1991
Earlier work this paper cites.
The continuum-armed bandit problem
Rajeev Agrawal · 1995
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 1995
Earlier work this paper cites.
Bandit problems with infinitely many arms
Donald A. Berry, Robert W. Chen, Alan Zame, David C. Heath, and Larry A. Shepp · 1997
Earlier work this paper cites.
Empirical support for winnow and weighted-majority based algorithms: Results on a calendar scheduling domain
Avrim Blum · 1997
Earlier work this paper cites.
How to use expert advice
Nicolò Cesa-Bianchi, Yoav Freund, David Haussler, David P. Helmbold, Robert E. Schapire, and Manfred K. Warmuth · 1997
Earlier work this paper cites.
Using and combining predictors that specialize
Yoav Freund, Robert E Schapire, Yoram Singer, and Manfred K Warmuth · 1997
Earlier work this paper cites.
A game of prediction with expert advice
V. Vovk · 1998
Earlier work this paper cites.
Deterministic Global Optimization: Theory, Algorithms and Applications
Christodoulos A. Floudas · 1999
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2000
Earlier work this paper cites.
A Metric for Distributions with Applications to Image Databases
Yossi Rubner, Carlo Tomasi, , and Leonidas J. Guibas · 2000
Earlier work this paper cites.
Lectures on analysis on metric spaces
J. Heinonen · 2001
Earlier work this paper cites.
Finding Nearest Neighbors in Growth-restricted Metrics
D.R. Karger and M. Ruhl · 2002
Earlier work this paper cites.
Online learning in online auctions
Avrim Blum, Vijay Kumar, Atri Rudra, and Felix Wu · 2003
Earlier work this paper cites.
Bounded geometries, fractals, and low–distortion embeddings
Anupam Gupta, Robert Krauthgamer, and James R. Lee · 2003
Earlier work this paper cites.
The value of knowing a demand curve: Bounds on regret for online posted-price auctions
Robert D. Kleinberg and Frank T. Leighton · 2003
Earlier work this paper cites.
Online linear optimization and adaptive routing
Baruch Awerbuch and Robert Kleinberg · 2004
Earlier work this paper cites.
Regret and convergence bounds for immediate-reward reinforcement learning with continuous action spaces
Eric Cope · 2004
Earlier work this paper cites.
Object location in realistic networks
Kirsten Hildrum, John Kubiatowicz, and Satish Rao · 2004
Earlier work this paper cites.
Nearly tight bounds for the continuum-armed bandit problem
Robert Kleinberg · 2004
Earlier work this paper cites.
Bypassing the embedding: Algorithms for low-dimensional metrics
Kunal Talwar · 2004
Earlier work this paper cites.
Name independent routing for growth bounded networks
Ittai Abraham and Dahlia Malkhi · 2005
Earlier work this paper cites.
On hierarchical routing in bounded-growth metrics
Hubert T-H. Chan, Anupam Gupta, Bruce M. Maggs, and Shuheng Zhou · 2005
Cited alongside, same era.
Online Convex Optimization in the Bandit Setting: Gradient Descent without a Gradient
Abraham Flaxman, Adam Kalai, and H. Brendan McMahan · 2005
Cited alongside, same era.
Triangulation and embedding using small sets of beacons
Jon Kleinberg, Aleksandrs Slivkins, and Tom Wexler · 2005
Cited alongside, same era.
Online Decision Problems with Large Strategy Sets
Robert Kleinberg · 2005
Cited alongside, same era.
Fast construction of nets in low dimensional metrics, and their applications
Manor Mendel and Sariel Har-Peled · 2005
Cited alongside, same era.
Probability and Computing: Randomized Algorithms and Probabilistic Analysis
Michael Mitzenmacher and Eli Upfal · 2005
An asymptotically optimal bandit algorithm for bounded support models
Junya Honda and Akimichi Takemura · 2010
Later among the works it cites.
Sharp dichotomies for regret minimization in metric spaces
Robert Kleinberg and Aleksandrs Slivkins · 2010
Later among the works it cites.
Showing Relevant Ads via Lipschitz Context Multi-Armed Bandits
Tyler Lu, Dávid Pál, and Martin Pál · 2010
Later among the works it cites.
Online Learning in Adversarial Lipschitz Environments
Odalric-Ambrym Maillard and Rémi Munos · 2010
Later among the works it cites.
Ranked bandits in metric spaces: Learning optimally diverse rankings over large document collections
Aleksandrs Slivkins, Filip Radlinski, and Sreenivas Gollapudi · 2010
Later among the works it cites.
Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design
Niranjan Srinivas, Andreas Krause, Sham Kakade, and Matthias Seeger · 2010
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distance estimation and object location via rings of neighbors
Aleksandrs Slivkins · 2005
Cited alongside, same era.
The Generic Chaining: Upper and Lower Bounds of Stochastic Processes
Michel Talagrand · 2005
Cited alongside, same era.
Bandit Problems
Dirk Bergemann and Juuso Välimäki · 2006
Cited alongside, same era.
Prediction, learning, and games
Nicolò Cesa-Bianchi and Gábor Lugosi · 2006
Cited alongside, same era.
Searching dynamic point sets in spaces with bounded doubling dimension
Richard Cole and Lee-Ad Gottlieb · 2006
Cited alongside, same era.
Bandit Based Monte-Carlo Planning
Levente Kocsis and Csaba Szepesvari · 2006
Cited alongside, same era.
Later among the works it cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Later among the works it cites.
Bandits, query learning, and the haystack dimension
Kareem Amin, Michael Kearns, and Umar Syed · 2011
Later among the works it cites.
The KL-UCB Algorithm for Bounded Stochastic Bandits and Beyond
Aurélien Garivier and Olivier Cappé · 2011
Later among the works it cites.
Multi-Armed Bandit Allocation Indices
John Gittins, Kevin Glazebrook, and Richard Weber · 2011
Later among the works it cites.
Contextual gaussian process bandit optimization
Andreas Krause and Cheng Soon Ong · 2011
Later among the works it cites.
Adaptive Bandits: Towards the best history-dependent strategy
Odalric-Ambrym Maillard and Rémi Munos · 2011
Later among the works it cites.
Optimistic optimization of a deterministic function without the knowledge of its smoothness
Rémi Munos · 2011
Later among the works it cites.
Multi-armed bandits on implicit metric spaces
Aleksandrs Slivkins · 2011
Later among the works it cites.
Contextual bandits with similarity information
Aleksandrs Slivkins · 2011
Later among the works it cites.
Dynamic pricing with limited supply
Moshe Babaioff, Shaddin Dughmi, Robert D. Kleinberg, and Aleksandrs Slivkins · 2012
Later among the works it cites.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Later among the works it cites.
Parallelizing exploration-exploitation tradeoffs with gaussian process bandit optimization
Thomas Desautels, Andreas Krause, and Joel Burdick · 2012
Later among the works it cites.
Truthful mechanisms with implicit payment computation
Moshe Babaioff, Robert Kleinberg, and Aleksandrs Slivkins · 2013
Closest in time.
Bandits with knapsacks
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Aleksandrs Slivkins · 2013
Closest in time.
Estimation of extreme values and associated level sets of a regression function via selective sampling
Stanislav Minsker · 2013
Closest in time.
Stochastic simultaneous optimistic optimization
Michal Valko, Alexandra Carpentier, and Rémi Munos · 2013
Closest in time.
Bandits with concave rewards and convex knapsacks
Shipra Agrawal and Nikhil R. Devanur · 2014
Closest in time.
Online stochastic optimization under correlated bandit feedback
Mohammad Gheshlaghi Azar, Alessandro Lazaric, and Emma Brunskill · 2014
Closest in time.
Adaptive contract design for crowdsourcing markets: Bandit algorithms for repeated principal-agent problems
Chien-Ju Ho, Aleksandrs Slivkins, and Jennifer Wortman Vaughan · 2014
Closest in time.
Lipschitz bandits: Regret lower bound and optimal algorithms
Stefan Magureanu, Richard Combes, and Alexandre Proutiere · 2014
Closest in time.
From bandits to monte-carlo tree search: The optimistic principle applied to optimization and planning
Rémi Munos · 2014
Closest in time.
Understanding Machine Learning : From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Closest in time.
Close the gaps: A learning-while-doing algorithm for single-product revenue management problems
Zizhuo Wang, Shiming Deng, and Yinyu Ye · 2014
Closest in time.
Adaptive-treed bandits
Adam Bull · 2015
Closest in time.
A near-optimal exploration-exploitation approach for assortment selection
Shipra Agrawal, Vashist Avadhanula, Vineet Goyal, and Assaf Zeevi · 2016
Closest in time.