Fetching the paper…
Reading the bibliography…
We present a new algorithm based on an gradient ascent for a general Active Exploration bandit problem in the fixed confidence setting.
An algorithm for quadratic programming
Marguerite Frank and Philip Wolfe · 1956
Earlier work this paper cites.
Sequential design of experiments
Herman Chernoff · 1959
Earlier work this paper cites.
The law of the iterated logarithm for empirical distribution
Helen Finkelstein et al · 1971
Earlier work this paper cites.
Pac bounds for multi-armed bandit and markov decision processes
Eyal Even-Dar, Shie Mannor, and Yishay Mansour · 2002
Earlier work this paper cites.
The sample complexity of exploration in the multi-armed bandit problem
Shie Mannor and John N Tsitsiklis · 2004
Earlier work this paper cites.
Active learning in multi-armed bandits
András Antos, Varun Grover, and Csaba Szepesvári · 2008
Earlier work this paper cites.
Self-Normalized Processes
Victor H Peña, Tze Leung Lai, and Qi-Man Shao · 2008
Earlier work this paper cites.
Best arm identification in multi-armed bandits
Jean-Yves Audibert and Sébastien Bubeck · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Introduction to online optimization
Sébastien Bubeck · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck, Nicolo Cesa-Bianchi, et al · 2012
Cited alongside, same era.
Online learning and online convex optimization
Shai Shalev-Shwartz et al · 2012
Cited alongside, same era.
Sharp finite-time iterated-logarithm martingale concentration
Akshay Balsubramani · 2014
Cited alongside, same era.
Best-arm identification in linear bandits
Marta Soare, Alessandro Lazaric, and Rémi Munos · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Cited alongside, same era.
Optimal best arm identification with fixed confidence
Aurélien Garivier and Emilie Kaufmann · 2016
Cited alongside, same era.
Fast rates for bandit optimization with upper-confidence frank-wolfe
Quentin Berthet and Vianney Perchet · 2017
Later among the works it cites.
Minimal exploration in structured stochastic bandits
Richard Combes, Stefan Magureanu, and Alexandre Proutiere · 2017
Later among the works it cites.
Thresholding bandit for dose-ranging: The impact of monotonicity
Aurélien Garivier, Pierre Ménard, and Laurent Rossi · 2017
Later among the works it cites.
The end of optimism? an asymptotic analysis of finite-armed linear bandits
Tor Lattimore and Csaba Szepesvari · 2017
Later among the works it cites.
The simulator: Understanding adaptive sampling in the moderate-confidence regime
Max Simchowitz, Kevin Jamieson, and Benjamin Recht · 2017
Later among the works it cites.
Explore first, exploit next: The true shape of regret in bandit problems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the complexity of best-arm identification in multi-armed bandit models
Emilie Kaufmann, Olivier Cappé, and Aurélien Garivier · 2016
Cited alongside, same era.
An optimal algorithm for the thresholding bandit problem
Andrea Locatelli, Maurilio Gutzeit, and Alexandra Carpentier · 2016
Cited alongside, same era.
Simple bayesian algorithms for best arm identification
Daniel Russo · 2016
Cited alongside, same era.
Aurélien Garivier, Pierre Ménard, and Gilles Stoltz · 2018
Later among the works it cites.
Mixture martingales revisited with applications to sequential tests and confidence intervals
Emilie Kaufmann and Wouter Koolen · 2018
Later among the works it cites.
Pure exploration with multiple correct answers
Rémy Degenne and Wouter M Koolen · 2019
Closest in time.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2019
Closest in time.