Hierarchical exploration for accelerating contextual bandits
Yue, Y · 1902
Earlier work this paper cites.
Risk, uncertainty and profit
Knight, F. H · 1921
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Sequential Analysis
Wald, A · 1947
Earlier work this paper cites.
Bayes and minimax solutions of sequential decision problems
Arrow, K. J · 1949
Earlier work this paper cites.
On information and sufficiency
Kullback, S · 1951
Earlier work this paper cites.
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
Chernoff, H · 1952
Earlier work this paper cites.
Adaptive Monte Carlo via bandit allocation
Neufeld, J · 1952
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Robbins, H · 1952
Earlier work this paper cites.
Continuous inspection schemes
Page, E · 1954
Earlier work this paper cites.
Factors influencing rate and extent of learning in the presence of misinformative feedback
Morin, R. E · 1955
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
Rosenblatt, F · 1958
Earlier work this paper cites.
Sequential medical trials
Armitage, P · 1960
Earlier work this paper cites.
Sequential medical trials
Anscombe, F · 1963
Earlier work this paper cites.
A model for selecting one of two medical treatments
Colton, T · 1963
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Hoeffding, W · 1963
Earlier work this paper cites.
Two-stage designs for clinical trials
Day, N · 1969
Earlier work this paper cites.
Inference about the change-point in a sequence of random variables
Hinkley, D. V · 1970
Earlier work this paper cites.
The randomized play-the-winner rule in medical trials
Wei, L · 1978
Earlier work this paper cites.
Bootstrap methods: another look at the Jackknife
Efron, B · 1979
Earlier work this paper cites.
Bandit processes and dynamic allocation indices
Gittins, J. C · 1979
Earlier work this paper cites.
A dynamic allocation index for the discounted multiarmed bandit problem
Gittins, J. C · 1979
Earlier work this paper cites.
A one-armed bandit problem with a concomitant variable
Woodroofe, M · 1979
Earlier work this paper cites.
Multi-armed bandits and the Gittins index
Whittle, P · 1980
Earlier work this paper cites.
Heuristics: intelligent search strategies for computer problem solving
Pearl, J · 1984
Earlier work this paper cites.
A theory of the learnable
Valiant, L. G · 1984
Earlier work this paper cites.
Disappointment in decision making under uncertainty
Bell, D. E · 1985
Earlier work this paper cites.
Bandit Problems: Sequential Allocation of Experiments
Berry, D · 1985
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L · 1985
Earlier work this paper cites.
The multi-armed bandit problem: decomposition and computation
Katehakis, M. N · 1987
Earlier work this paper cites.
Detecting changes in signals and systems: a survey
Basseville, M · 1988
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Statistical approaches to interim monitoring of medical trials: A review and commentary
Jennison, C · 1990
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S · 1990
Earlier work this paper cites.
On the Gittins index for multiarmed bandits
Weber, R · 1992
Earlier work this paper cites.
On adjusting p-values for multiplicity
Westfall, P. H · 1993
Earlier work this paper cites.
A course in game theory
Osborne, M. J · 1994
Earlier work this paper cites.
A short proof of the Gittins index theorem
Tsitsiklis, J. N · 1994
Earlier work this paper cites.
Sample mean based index policies with O(log n) regret for the multi-armed bandit problem
Agrawal, R · 1995
Earlier work this paper cites.
Decision-making under uncertainty: Capturing dynamic brand choice processes in turbulent consumer goods markets
Erdem, T · 1996
Earlier work this paper cites.
Multiple Comparisons: Theory and Methods
Hsu, J · 1996
Earlier work this paper cites.
Coping with uncertainty: A naturalistic decision-making analysis
Lipshitz, R · 1997
Earlier work this paper cites.
On-line learning with malicious noise and the closure algorithm
Auer, P · 1998
Earlier work this paper cites.
Finite-time regret bounds for the multiarmed bandit problem
Cesa-Bianchi, N · 1998
Earlier work this paper cites.
Randomized play-the-winner clinical trials: review and recommendations
Rosenberger, W. F · 1999
Earlier work this paper cites.
Learning and planning in structured worlds
Dearden, R. W · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P · 2002
Earlier work this paper cites.