Fetching the paper…
Reading the bibliography…
We revisit lower bounds on the regret in the case of multi-armed bandit problems.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. 1933 · 1933
Earlier work this paper cites.
A general class of coefficients of divergence of one distribution from another
Ali, S. M., S. D. Silvey. 1966 · 1966
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L., H. Robbins. 1985 · 1985
Earlier work this paper cites.
Optimal adaptive policies for sequential allocation problems
Burnetas, A.N., M.N. Katehakis. 1996 · 1996
Earlier work this paper cites.
Theory of Point Estimation
Lehmann, E.L., G. Casella. 1998 · 1998
Earlier work this paper cites.
Minimax lower bounds for the two-armed bandit problem
Kulkarni, S., G. Lugosi. 2000 · 2000
Earlier work this paper cites.
The sample complexity of exploration in the multi-armed bandit problem
Mannor, S., J.N. Tsitsiklis. 2004 · 2004
Earlier work this paper cites.
Prediction, Learning, and Games
Cesa-Bianchi, N., G. Lugosi. 2006 · 2006
Cited alongside, same era.
The exponential complexity of satisfiability problems
Calabro, Chris. 2009 · 2009
Cited alongside, same era.
UCB revisited: Improved regret bounds for the stochastic multi-armed bandit problem
Auer, P., R. Ortner. 2010 · 2010
Cited alongside, same era.
Bandits games and clustering foundations
Bubeck, S. 2010 · 2010
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S., N. Cesa-Bianchi. 2012 · 2012
Cited alongside, same era.
Kullback-Leibler upper confidence bounds for optimal sequential allocation
Cappé, O., A. Garivier, O.-A. Maillard, R. Munos, G. Stoltz. 2013 · 2013
Cited alongside, same era.
Asymptotically optimal sequential experimentation under generalized ranking
Cowan, W., M.N. Katehakis. 2015 · 2015
Later among the works it cites.
Online learning and game theory. a quick overview with recent results and applications
Faure, M., P. Gaillard, B. Gaujal, V. Perchet. 2015 · 2015
Later among the works it cites.
Non-asymptotic analysis of a new bandit algorithm for semi-bounded rewards
Honda, J., A. Takemura. 2015 · 2015
Later among the works it cites.
Online advertisements and multi-armed bandits
Jiang, C. 2015 · 2015
Later among the works it cites.
Online learning with Gaussian payoffs and side observations
Wu, Y., A. György, C. Szepesvari. 2015 · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Combes, R., A. Proutière. 2014 · 2014
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Auer, P., N. Cesa-Bianchi, P. Fischer. 2002a
Cited in the paper.
The nonstochastic multiarmed bandit problem
Auer, P., N. Cesa-Bianchi, Y. Freund, R.E. Schapire. 2002b
Cited in the paper.
Bounded regret in stochastic multi-armed bandits
Bubeck, S., V. Perchet, P. Rigollet. 2013a
Cited in the paper.
Erratum to [ 7 ]
Bubeck, S., V. Perchet, P. Rigollet. 2013b
Cited in the paper.
Garivier, A., E. Kaufmann, T. Lattimore. 2016 · 2016
Closest in time.
On the complexity of best arm identification in multi-armed bandit models
Kaufmann, E., O. Cappé, A. Garivier. 2016 · 2016
Closest in time.