Fetching the paper…
Reading the bibliography…
We study the best-arm identification problem with fixed confidence when contextual (covariate) information is available in stochastic bandits.
Neyman, J. (1923), “Sur les applications de la theorie des probabilites aux experiences agricoles: Essai des principes,” Statistical Science
1923
Earlier work this paper cites.
Thompson, W. R. (1933), “On the likelihood that one unknown probability exceeds another in view of the evidence of two samples,” Biometrika
1933
Earlier work this paper cites.
Robbins, H. (1952), “Some aspects of the sequential design of experiments,” Bulletin of the American Mathematical Society
1952
Earlier work this paper cites.
Chernoff, H. (1959), “Sequential Design of Experiments,” The Annals of Mathematical Statistics
1959
Earlier work this paper cites.
Paulson, E. (1964), “A Sequential Procedure for Selecting the Population with the Largest Mean from k k Normal Populations,” The Annals of Mathematical Statistics
1964
Earlier work this paper cites.
Bechhofer, R., Kiefer, J., and Sobel, M. (1968), Sequential Identification and Ranking Procedures: With Special Reference to Koopman-Darmois Populations
1968
Earlier work this paper cites.
Hogan, W. W. (1973), “Point-to-set maps in mathematical programming,” SIAM review
1973
Earlier work this paper cites.
Rubin, D. B. (1974), “Estimating causal effects of treatments in randomized and nonrandomized studies,” Journal of Educational Psychology
1974
Earlier work this paper cites.
Lai, T. L. and Robbins, H. (1985), “Asymptotically efficient adaptive allocation rules,” Advances in Applied Mathematics
1985
Earlier work this paper cites.
Mannor, S. and Tsitsiklis, J. N. (2004), “The sample complexity of exploration in the multi-armed bandit problem,” Journal of Machine Learning Research
2004
Earlier work this paper cites.
Even-Dar, E., Mannor, S., Mansour, Y., and Mahadevan, S. (2006), “Action Elimination and Stopping Conditions for the Multi-Armed Bandit and Reinforcement Learning Problems.” Journal of Machine Learning Research
2006
Earlier work this paper cites.
Antos, A., Grover, V., and Szepesvári, C. (2008), “Active Learning in Multi-armed Bandits,” in Algorithmic Learning Theory
2008
Earlier work this paper cites.
van der Laan, M. J. (2008), “The Construction and Analysis of Adaptive Group Sequential Designs,”
2008
Earlier work this paper cites.
Bubeck, S., Munos, R., and Stoltz, G. (2011), “Pure exploration in finitely-armed and continuous-armed bandits,” Theoretical Computer Science
2011
Earlier work this paper cites.
Hahn, J., Hirano, K., and Karlan, D. (2011), “Adaptive experimental design using the propensity score,” Journal of Business and Economic Statistics
2011
Cited alongside, same era.
Gabillon, V., Ghavamzadeh, M., and Lazaric, A. (2012), “Best Arm Identification: A Unified Approach to Fixed Budget and Fixed Confidence,” in Advances in Neural Information Processing Systems
2012
Cited alongside, same era.
Cappé, O., Garivier, A., Maillard, O.-A., Munos, R., and Stoltz, G. (2013), “Kullback–Leibler upper confidence bounds for optimal sequential allocation,” The Annals of Statistics
2013
Cited alongside, same era.
Chiu, S., Stoyan, D., Kendall, W., and Mecke, J. (2013), Stochastic Geometry and Its Applications
2013
Cited alongside, same era.
Karnin, Z., Koren, T., and Somekh, O. (2013), “Almost optimal exploration in multi-armed bandits,” in International Conference on Machine Learning
Deshmukh, A. A., Sharma, S., Cutler, J. W., Moldwin, M., and Scott, C. (2018), “Simple Regret Minimization for Contextual Bandits,”
2018
Later among the works it cites.
Guan, M. Y. and Jiang, H. (2018), “Nonparametric Stochastic Contextual Bandits,” in AAAI Conference on Artificial Intelligence
2018
Later among the works it cites.
Tabord-Meehan, M. (2018), “Stratification Trees for Adaptive Randomization in Randomized Controlled Trials,”
2018
Later among the works it cites.
Tao, C., Blanco, S., and Zhou, Y. (2018), “Best Arm Identification in Linear Bandits with Linear Dimension Dependency,” in International Conference on Machine Learning
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
Jamieson, K., Malloy, M., Nowak, R., and Bubeck, S. (2014), “lil’ UCB : An Optimal Exploration Algorithm for Multi-Armed Bandits,” in Conference on Learning Theory
2014
Cited alongside, same era.
Karlan, D. and Wood, D. H. (2014), “The Effect of Effectiveness: Donor Response to Aid Effectiveness in a Direct Mail Fundraising Experiment,” Working paper, National Bureau of Economic Research
2014
Cited alongside, same era.
Soare, M., Lazaric, A., and Munos, R. (2014), “Best-Arm Identification in Linear Bandits,” in Advances in Neural Information Processing Systems
2014
Cited alongside, same era.
Imbens, G. W. and Rubin, D. B. (2015), Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction
2015
Cited alongside, same era.
Tekin, C. and van der Schaar, M. (2015), “RELEAF: An Algorithm for Learning and Exploiting Relevance,” IEEE Journal of Selected Topics in Signal Processing
2015
Cited alongside, same era.
Athey, S. and Imbens, G. (2016), “The Econometrics of Randomized Experiments,”
2016
Cited alongside, same era.
Garivier, A. and Kaufmann, E. (2016), “Optimal Best Arm Identification with Fixed Confidence,” in Conference on Learning Theory
2016
Cited alongside, same era.
2018
Later among the works it cites.
Degenne, R., Koolen, W. M., and Ménard, P. (2019), “Non-Asymptotic Pure Exploration by Solving Games,” in Advances in Neural Information Processing Systems
2019
Later among the works it cites.
Fiez, T., Jain, L., Jamieson, K. G., and Ratliff, L. (2019), “Sequential Experimental Design for Transductive Linear Bandits,” in Advances in Neural Information Processing Systems
2019
Later among the works it cites.
Garivier, A., Ménard, P., and Stoltz, G. (2019), “Explore first, exploit next: The true shape of regret in bandit problems,” Mathematics of Operations Research
2019
Later among the works it cites.
Juneja, S. and Krishnasamy, S. (2019), “Sample complexity of partition identification using multi-armed bandits,” in Conference on Learning Theory
2019
Later among the works it cites.
Jedra, Y. and Proutiere, A. (2020), “Optimal Best-arm Identification in Linear Bandits,” Advances in Neural Information Processing Systems
2020
Later among the works it cites.
Kato, M., Ishihara, T., Honda, J., and Narita, Y. (2020), “Adaptive Experimental Design for Efficient Treatment Effect Estimation,”
2020
Later among the works it cites.
Kaufmann, E. and Koolen, W. M. (2021), “Mixture Martingales Revisited with Applications to Sequential Tests and Confidence Intervals,” Journal of Machine Learning Research
2021
Closest in time.
Russac, Y., Katsimerou, C., Bohle, D., Cappé, O., Garivier, A., and Koolen, W. M. (2021), “A/B/n Testing with Control in the Presence of Subpopulations,” in Advances in Neural Information Processing Systems
2021
Closest in time.
Qin, C. and Russo, D. (2022), “Adaptivity and Confounding in Multi-Armed Bandit Experiments,”
2022
Closest in time.