Fetching the paper…
Reading the bibliography…
We propose information-directed sampling -- a new approach to online optimization problems in which a decision-maker must balance between exploration and exploitation while learning from partial feedback.
On a measure of the information provided by an experiment
D. V. Lindley · 1956
Earlier work this paper cites.
A new method of locating the maximum point of an arbitrary multipeak curve in the presence of noise
H.J. Kushner · 1964
Earlier work this paper cites.
The application of Bayesian methods for seeking the extremum
J. Mockus, V. Tiesis, and A. Zilinskas · 1978
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T.L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Adaptive treatment allocation and the multi-armed bandit problem
T.L. Lai · 1987
Earlier work this paper cites.
Bayesian experimental design: A review
K. Chaloner, I. Verdinelli, et al · 1995
Earlier work this paper cites.
Asymptotically efficient adaptive choice of control laws in controlled Markov chains
T.L. Graves and T.L. Lai · 1997
Earlier work this paper cites.
Discrete prediction games with arbitrary feedback and loss
A. Piccolboni and C. Schindelhauer · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Convex optimization
S.P. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Bandit based Monte-Carlo planning
L. Kocsis and Cs. Szepesvári · 2006
Earlier work this paper cites.
The price of bandit information for online optimization
V. Dani, S.M. Kakade, and T.P. Hayes · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
V. Dani, T.P. Hayes, and S.M. Kakade · 2008
Earlier work this paper cites.
A knowledge-gradient policy for sequential information collection
P.I. Frazier, W.B. Powell, and S. Dayanik · 2008
Earlier work this paper cites.
Multi-armed bandits in metric spaces
R. Kleinberg, A. Slivkins, and E. Upfal · 2008
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
J.-Y. Audibert and S. Bubeck · 2009
Earlier work this paper cites.
A tutorial on Bayesian optimization of expensive cost functions, with application to active user modeling and hierarchical reinforcement learning
E. Brochu, V.M. Cora, and N. de Freitas · 2009
Earlier work this paper cites.
An informational approach to the global optimization of expensive-to-evaluate functions
J. Villemonteix, E. Vazquez, and E. Walter · 2009
Earlier work this paper cites.
Parametric bandits: The generalized linear case
S. Filippi, O. Cappé, A. Garivier, and C. Szepesvári · 2010
Earlier work this paper cites.
Paradoxes in learning and the marginal value of information
P.I. Frazier and W.B. Powell · 2010
Earlier work this paper cites.
Near-optimal bayesian active learning with noisy observations
D. Golovin, A. Krause, and D. Ray · 2010
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Cited alongside, same era.
Linearly parameterized bandits
P. Rusmevichientong and J.N. Tsitsiklis · 2010
Cited alongside, same era.
Dynamic assortment optimization with a multinomial logit choice model and capacity constraint
P. Rusmevichientong, Z.-J. M. Shen, and D.B. Shmoys · 2010
Cited alongside, same era.
On the robustness of a one-period look-ahead policy in multi-armed bandit problems
I. Ryzhov, P. Frazier, and W. Powell · 2010
Cited alongside, same era.
A modern Bayesian look at the multi-armed bandit
S.L. Scott · 2010
Cited alongside, same era.
Thompson sampling: an asymptotically optimal finite time analysis
E. Kaufmann, N. Korda, and R. Munos · 2012
Later among the works it cites.
Optimal learning , volume 841
W.B. Powell and I.O. Ryzhov · 2012
Later among the works it cites.
The knowledge gradient algorithm for a general class of online learning problems
I.O. Ryzhov, W.B. Powell, and P.I. Frazier · 2012
Later among the works it cites.
Information-theoretic regret bounds for Gaussian process optimization in the bandit setting
N. Srinivas, A. Krause, S.M. Kakade, and M. Seeger · 2012
Later among the works it cites.
Regret in online combinatorial optimization
J.-Y. Audibert, S. Bubeck, and G. Lugosi · 2013
Later among the works it cites.
Bayesian mixture modelling and inference based Thompson sampling in Monte-Carlo tree search
A. Bai, F. Wu, and X. Chen · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Cited alongside, same era.
X-armed bandits
S. Bubeck, R. Munos, G. Stoltz, and C. Szepesvári · 2011
Cited alongside, same era.
An empirical evaluation of Thompson sampling
O. Chapelle and L. Li · 2011
Cited alongside, same era.
Multi-Armed Bandit Allocation Indices
J. Gittins, K. Glazebrook, and R. Weber · 2011
Cited alongside, same era.
Adaptive submodularity: Theory and applications in active learning and stochastic optimization
D. Golovin and A. Krause · 2011
Cited alongside, same era.
Entropy and information theory
R.M. Gray · 2011
Cited alongside, same era.
Later among the works it cites.
Kullback-Leibler upper confidence bounds for optimal sequential allocation
O. Cappé, A. Garivier, O.-A. Maillard, R. Munos, and G. Stoltz · 2013
Later among the works it cites.
(More) efficient reinforcement learning via posterior sampling
I. Osband, D. Russo, and B. Van Roy · 2013
Later among the works it cites.
Eluder dimension and the sample complexity of optimistic exploration
D. Russo and B. Van Roy · 2013
Later among the works it cites.
Optimal dynamic assortment planning with demand learning
D. Sauré and A. Zeevi · 2013
Later among the works it cites.
Bisection search with noisy responses
R. Waeber, P.I. Frazier, and S.G. Henderson · 2013
Later among the works it cites.
Partial monitoring-classification, regret bounds, and algorithms
G. Bartók, D. P. Foster, D. Pál, A. Rakhlin, and C. Szepesvári · 2014
Closest in time.
Gaussian process optimization with mutual information
E. Contal, V. Perchet, and N. Vayatis: · 2014
Closest in time.
Thompson sampling for complex online problems
A. Gopalan, S. Mannor, and Y. Mansour · 2014
Closest in time.
Predictive entropy search for efficient global optimization of black-box functions
J. M. Hernández-Lobato, M. W. Hoffman, and Z. Ghahramani · 2014
Closest in time.
Bandit convex optimization: T \sqrt{T} regret in one dimension
S. Bubeck, O. Dekel, T. Koren, and Y. Peres · 2015
Closest in time.
Refined knowledge-gradient policy for learning probabilities
B. Kamiński · 2015
Closest in time.
Multi-scale exploration of convex functions and bandit convex optimization
S. Bubeck and R. Eldan · 2016
Closest in time.
An information-theoretic analysis of Thompson sampling
D. Russo and B. Van Roy · 2016
Closest in time.