Fetching the paper…
Reading the bibliography…
The analysis of online least squares estimation is at the heart of many stochastic sequential decision making problems.
Some aspects of the sequential design of experiments
H. Robbins · 1952
Earlier work this paper cites.
Boundary crossing probabilities for the Wiener process and sample sums
H. Robbins and D. Siegmund · 1970
Earlier work this paper cites.
On tail probabilities for martingales
D.A. Freedman · 1975
Earlier work this paper cites.
Strong consistency of least squares estimates in multiple regression
T.L. Lai, H. Robbins, and C.Z. Wei · 1979
Earlier work this paper cites.
Least squares estimates in stochastic regression models with applications to identification and control of dynamic systems
T.L. Lai and C.Z. Wei · 1982
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Matrix Perturbation Theory
G.W. Stewart and Ji-guang Sun · 1990
Cited alongside, same era.
Finite time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Cited alongside, same era.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2003
Cited alongside, same era.
Self-normalized processes: exponential inequalities, moment bounds and iterated logarithm laws
V.H. de la Peña, M.J. Klass, and T.L. Lai · 2004
Cited alongside, same era.
Prediction, Learning, and Games
N. Cesa-Bianchi and G. Lugosi · 2006
Cited alongside, same era.
Online optimization in X-armed bandits
S. Bubeck, R. Munos, G. Stoltz, and Cs. Szepesvári · 2008
Cited alongside, same era.
Stochastic linear optimization under bandit feedback
V. Dani, T.P. Hayes, and S.M. Kakade · 2008
Later among the works it cites.
On upper-confidence bound policies for non-stationary bandit problems
A Garivier and E Moulines · 2008
Later among the works it cites.
Self-normalized processes: Limit theory and Statistical Applications
V.H. de la Peña, T.L. Lai, and Q.-M. Shao · 2009
Later among the works it cites.
Active learning in heteroscedastic noise
A. Antos, V. Grover, and Cs. Szepesvári · 2010
Later among the works it cites.
Linearly parameterized bandits
P. Rusmevichientong and J.N. Tsitsiklis · 2010
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…