Fetching the paper…
Reading the bibliography…
We consider a generalization of stochastic bandits where the set of arms, $\cX$, is allowed to be a generic measurable space and the mean-payoff function is "locally Lipschitz" with respect to a dissimilarity function that is known to the decision maker.
The continuum-armed bandit problem
R. Agrawal · 1951
Earlier work this paper cites.
Some aspects of the sequential design of experiments
H. Robbins · 1952
Earlier work this paper cites.
Stochastic Processes
J. L. Doob · 1953
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
W. Hoeffding · 1963
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Multi-armed Bandit Allocation Indices
J. C. Gittins · 1989
Earlier work this paper cites.
Nearly tight bounds for the continuum-armed bandit problem
R. Kleinberg · 2004
Earlier work this paper cites.
Prediction, Learning, and Games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Modification of UCT with patterns in Monte-Carlo go
S. Gelly, Y. Wang, R. Munos, and O. Teytaud · 2006
Cited alongside, same era.
Bandit based Monte-carlo planning
L. Kocsis and Cs. Szepesvari · 2006
Cited alongside, same era.
Improved rates for the stochastic continuum-armed bandit problem
P. Auer, R. Ortner, and C. Szepesvári · 2007
Cited alongside, same era.
Bandit algorithms for tree search
P.-A. Coquelin and R. Munos · 2007
Cited alongside, same era.
Combining online and offline knowledge in UCT
S. Gelly and D. Silver · 2007
Cited alongside, same era.
How powerful can any regression learning procedure be?
Y. Yang · 2007
Cited alongside, same era.
Simulation-based approach to general game playing
H. Finnsson and Y. Bjornsson · 2008
Later among the works it cites.
Achieving master level play in 9 × \times 9 computer go
S. Gelly and D. Silver · 2008
Later among the works it cites.
Addressing NP-complete puzzles with Monte-Carlo methods
M.P.D. Schadd, M.H.M. Winands, H.J. van den Herik, and H. Aldewereld · 2008
Later among the works it cites.
Online optimization in x-armed bandits
S. Bubeck, R. Munos, G. Stoltz, and Cs. Szepesvari · 2009
Later among the works it cites.
Regret and convergence bounds for immediate-reward reinforcement learning with continuous action spaces
E. Cope · 2009
Later among the works it cites.
Open loop optimistic planning
S. Bubeck and R. Munos · 2010
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Competing in the dark: an efficient algorithm for bandit linear optimization
J. Abernethy, E. Hazan, and A. Rakhlin · 2008
Cited alongside, same era.
Progressive strategies for Monte-Carlo tree search
G.M.J. Chaslot, M.H.M. Winands, H. Herik, J. Uiterwijk, and B. Bouzy · 2008
Cited alongside, same era.
Sample mean based index policies with o(log n) regret for the multi-armed bandit problem
R. Agrawal
Cited in the paper.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer
Cited in the paper.
The non-stochastic multi-armed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. Schapire
Cited in the paper.
Multi-armed bandits in metric spaces
R. Kleinberg, A. Slivkins, and E. Upfal
Cited in the paper.
Pure exploration in multi-armed bandits problems
S. Bubeck, R. Munos, and G. Stoltz · 2010
Closest in time.