Fetching the paper…
Reading the bibliography…
We consider bandit problems involving a large (possibly infinite) collection of arms, in which the expected reward of each arm is a linear function of an $r$-dimensional random vector $\mathbf{Z} \in \mathbb{R}^r$, where $r \geq 2$.
Self-normalized processes: exponential inequalities, moment bounds and iterated logarithm laws
De la Peña, V. H., M. J. Klass, and T. L. Lai. 2004 · 1933
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. 1933 · 1933
Earlier work this paper cites.
Adjustment of an inverse matrix corresponding to a change in one element of a given matrix
Sherman, J., and W. J. Morrison. 1950 · 1950
Earlier work this paper cites.
A stochastic approximation method
Robbins, H., and S. Monro. 1951 · 1951
Earlier work this paper cites.
Stochastic estimation of the maximum of a regression function
Kiefer, J., and J. Wolfowitz. 1952 · 1952
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Robbins, H. 1952 · 1952
Earlier work this paper cites.
Multidimensional Stochastic Approximation Methods
Blum, J. R. 1954 · 1954
Earlier work this paper cites.
Contributions to the “two-armed bandit” problem
Feldman, D. 1962 · 1962
Earlier work this paper cites.
Bandit Problems: Sequential Allocation of Experiments
Berry, D., and B. Fristedt. 1985 · 1985
Earlier work this paper cites.
Further contributions to the “two-armed bandit” problem
Keener, R. 1985 · 1985
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L., and H. Robbins. 1985 · 1985
Earlier work this paper cites.
Adaptive treatment allocation and the multi-armed bandit problem
Lai, T. L. 1987 · 1987
Cited alongside, same era.
Asymptotically efficient adaptive allocation schemes for controlled I.I.D. processes: finite parameter space
Agrawal, R., D. Teneketzis, and V. Anantharam. 1989 · 1989
Cited alongside, same era.
Sequential Control With Incomplete Information
Pressman, E. L., and I. N. Sonin. 1990 · 1990
Cited alongside, same era.
Sample mean based index policies with O(log n) regret for the multi-armed bandit problem
Agrawal, R. 1995 · 1995
Cited alongside, same era.
Dynamic Programming and Optimal Control
Bertsekas, D. 1995 · 1995
Cited alongside, same era.
Response surface bandits
Ginebra, J., and M. K. Clayton. 1995 · 1995
Cited alongside, same era.
Uniform Central Limit Theorems
Dudley, R. M. 1999 · 1999
Later among the works it cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P. 2002 · 2002
Later among the works it cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., N. Cesa-Bianchi, and P. Fischer. 2002 · 2002
Later among the works it cites.
Stochastic Approximation (Invited Paper)
Lai, T. 2003 · 2003
Later among the works it cites.
Positive Definite Matrices
Bhatia, R. 2007 · 2007
Later among the works it cites.
Multi-armed bandit problems with dependent arms
Pandey, S., D. Chakrabarti, and D. Agrawal. 2007 · 2007
Later among the works it cites.
Stochastic linear optimization under bandit feedback
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neuro-Dynamic Programming
Bertsekas, D., and J. N. Tsitsiklis. 1996 · 1996
Cited alongside, same era.
Strongly convex analysis
Polovinkin, E. S. 1996 · 1996
Cited alongside, same era.
Introduction to Linear Optimization
Bertsimas, D., and J. N. Tsitsiklis. 1997 · 1997
Cited alongside, same era.
A new positive definite geometric mean of two positive defintie matrices
Fiedler, M., and V. Pták. 1997 · 1997
Cited alongside, same era.
Associative reinforcement learning using linear probabilistic concepts
Abe, N., and P. M. Long. 1999 · 1999
Cited alongside, same era.
Stochastic linear optimization under bandit feedback
Dani, V., T. P. Hayes, and S. M. Kakade. December 2008b
Cited in the paper.
Dani, V., T. P. Hayes, and S. M. Kakade. 2008a · 2008
Closest in time.
Performance limitations in bandit problems with side observations
Goldenshluger, A., and A. Zeevi. 2008 · 2008
Closest in time.
General bounds and finite-time performance improvement for the Kiefer-Wolfowitz stochastic approximation algorithm
Cicek, D., M. Broadie, and A. Zeevi. 2009 · 2009
Closest in time.
Woodroofe’s one-armed bandit problem revisited
Goldenshluger, A., and A. Zeevi. 2009 · 2009
Closest in time.
A structured multiarmed bandit problem and the greedy policy
Mersereau, A. J., P. Rusmevichientong, and J. N. Tsitsiklis. 2009 · 2009
Closest in time.