Fetching the paper…
Reading the bibliography…
In the stochastic bandit problem, the goal is to maximize an unknown function via a sequence of noisy evaluations.
On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples
William R. Thompson · 1933
Earlier work this paper cites.
IV. on least squares and linear combination of observations
Alexander C Aitken · 1936
Earlier work this paper cites.
Convergence of Random Processes and Limit Theorems in Probability Theory
Y. Prokhorov · 1956
Earlier work this paper cites.
On tail probabilities for martingales
David A Freedman · 1975
Earlier work this paper cites.
Optimal adaptive policies for sequential allocation problems
Apostolos N Burnetas and Michael N Katehakis · 1996
Earlier work this paper cites.
Regression with Input-Dependent Noise: A Gaussian Process Treatment
Paul Goldberg, Christopher K. I. Williams, and Christopher M. Bishop · 1998
Earlier work this paper cites.
A better bound on the variance
Rajendra Bhatia and Chandler Davis · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Gaussian Processes for Machine Learning , volume 1
Carl Edward Rasmussen and Christopher KI Williams · 2006
Earlier work this paper cites.
Most Likely Heteroscedastic Gaussian Process Regression
Kristian Kersting, Christian Plagemann, Patrick Pfaff, and Wolfram Burgard · 2007
Earlier work this paper cites.
Stochastic Linear Optimization under Bandit Feedback
Varsha Dani, Thomas P. Hayes, and Sham M. Kakade · 2008
Earlier work this paper cites.
On the Generalization Ability of Online Strongly Convex Programming Algorithms
Sham M Kakade and Ambuj Tewari · 2009
Cited alongside, same era.
Active learning in heteroscedastic noise
András Antos, Varun Grover, and Csaba Szepesvári · 2010
Cited alongside, same era.
Probability: Theory and Examples
Rick Durrett · 2010
Cited alongside, same era.
Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design
Niranjan Srinivas, Andreas Krause, Matthias Seeger, and Sham M Kakade · 2010
Cited alongside, same era.
Improved Algorithms for Linear Stochastic Bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Cited alongside, same era.
Online Learning for Linearly Parametrized Control Problems
Yasin Abbasi-Yadkori · 2012
Cited alongside, same era.
Wesley Cowan, Junya Honda, and Michael N. Katehakis · 2015
Later among the works it cites.
Exponential inequalities for martingales with applications
Xiequan Fan, Ion Grama, and Quansheng Liu · 2015
Later among the works it cites.
Linear Multi-Resource Allocation with Semi-Bandit Feedback
Tor Lattimore, Koby Crammer, and Csaba Szepesvari · 2015
Later among the works it cites.
An information-theoretic analysis of Thompson sampling
Daniel Russo and Benjamin Van Roy · 2016
Later among the works it cites.
Linear Thompson Sampling Revisited
Marc Abeille and Alessandro Lazaric · 2017
Later among the works it cites.
Active Heteroscedastic Regression
Kamalika Chaudhuri, Prateek Jain, and Nagarajan Natarajan · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Cited alongside, same era.
Elements of Information Theory
Thomas M Cover and Joy A Thomas · 2012
Cited alongside, same era.
Thompson Sampling for Contextual Bandits with Linear Payoffs
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
Heteroscedastic treed bayesian optimisation
John-Alexander M Assael, Ziyu Wang, Bobak Shahriari, and Nando de Freitas · 2014
Cited alongside, same era.
Learning to Optimize via Information-Directed Sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Later among the works it cites.
On Kernelized Multi-armed Bandits
Sayak Ray Chowdhury and Aditya Gopalan · 2017
Later among the works it cites.
A Scale Free Algorithm for Stochastic Bandits with Bounded Kurtosis
Tor Lattimore · 2017
Later among the works it cites.
The End of Optimism? An Asymptotic Analysis of Finite-Armed Linear Bandits
Tor Lattimore and Csaba Szepesvari · 2017
Later among the works it cites.
Annie Marsden and Sergio Bacallado · 2017
Later among the works it cites.
Max-value Entropy Search for Efficient Bayesian Optimization
Zi Wang and Stefanie Jegelka · 2017
Later among the works it cites.