Fetching the paper…
Reading the bibliography…
We study the problem of dynamic batch learning in high-dimensional sparse linear contextual bandits, where a decision maker, under a given maximum-number-of-batch constraint and only able to observe rewards at the end of each batch, can dynamically decide how many individuals to include in the next batch (at the end of the current batch) and what personalized action-selection scheme to adopt within each batch.
Some aspects of the sequential design of experiments
Robbins, H. (1952) · 1952
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P. (2002) · 2002
Earlier work this paper cites.
Sequential batch learning in finite-action linear contextual bandits
Han, Y., Zhou, Z., Zhou, Z., Blanchet, J., Glynn, P. W., and Ye, Y. (2020) · 2004
Earlier work this paper cites.
Elements of Information Theory
Cover, T. M. and Thomas, J. A. (2006) · 2006
Earlier work this paper cites.
A learning approach for interactive marketing to a customer segment
Bertsimas, D. and Mersereau, A. J. (2007) · 2007
Earlier work this paper cites.
Adaptive design methods in clinical trials–a review
Chow, S.-C. and Chang, M. (2008) · 2008
Earlier work this paper cites.
Challenges and opportunities in high-dimensional choice data analyses
Naik, P., Wedel, M., Bacon, L., Bodapati, A., Bradlow, E., Kamakura, W., Kreulen, J., Lenk, P., Madigan, D. M., and Montgomery, A. (2008) · 2008
Earlier work this paper cites.
Simultaneous analysis of lasso and dantzig selector
Bickel, P. J., Ritov, Y., Tsybakov, A. B., et al. (2009) · 2009
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Filippi, S., Cappe, O., Garivier, A., and Szepesvári, C. (2010) · 2010
Earlier work this paper cites.
Nonparametric bandits with covariates
Rigollet, P. and Zeevi, A. (2010) · 2010
Earlier work this paper cites.
High dimensional sparse econometric models: An introduction
Belloni, A. and Chernozhukov, V. (2011) · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Chu, W., Li, L., Reyzin, L., and Schapire, R. (2011) · 2011
Earlier work this paper cites.
The battle trial: personalizing therapy for lung cancer
Kim, E. S., Herbst, R. S., Wistuba, I. I., Lee, J. J., Blumenschein, G. R., Tsao, A., Stewart, D. J., Hicks, M. E., Erasmus, J., Gupta, S., et al. (2011) · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S., Cesa-Bianchi, N., et al. (2012) · 2012
Earlier work this paper cites.
Bandit theory meets compressed sensing for high dimensional stochastic linear bandit
Carpentier, A. and Munos, R. (2012) · 2012
Earlier work this paper cites.
Online learning for linearly parametrized control problems
Abbasi-Yadkori, Y. (2013) · 2013
Earlier work this paper cites.
A linear response bandit problem
Goldenshluger, A. and Zeevi, A. (2013) · 2013
Earlier work this paper cites.
Data-driven decisions for reducing readmissions for heart failure: General methodology and case study
Bayati, M., Braverman, M., Gillam, M., Mack, K. M., Ruiz, G., Smith, M. S., and Horvitz, E. (2014) · 2014
Cited alongside, same era.
Inference on treatment effects after selection among high-dimensional controls
Belloni, A., Chernozhukov, V., and Hansen, C. (2014) · 2014
Cited alongside, same era.
Doubly robust learning for estimating individualized treatment with censored data
Zhao, Y.-Q., Zeng, D., Laber, E. B., Song, R., Yuan, M., and Kosorok, M. R. (2014) · 2014
Cited alongside, same era.
Statistical Learning with Sparsity: The Lasso and Generalizations
Hastie, T., Tibshirani, R., and Wainwright, M. (2015) · 2015
Cited alongside, same era.
Population-level prediction of type 2 diabetes from claims data and analysis of risk factors
Razavian, N., Blecker, S., Schmidt, A. M., Smith-McLallen, A., Nigam, S., and Sontag, D. (2015) · 2015
Cited alongside, same era.
Confounding-robust policy improvement
Kallus, N. and Zhou, A. (2018) · 2018
Later among the works it cites.
Who should be treated? empirical welfare maximization methods for treatment choice
Kitagawa, T. and Tetenov, A. (2018) · 2018
Later among the works it cites.
Bandit algorithms
Lattimore, T. and Szepesvári, C. (2018) · 2018
Later among the works it cites.
Adaptive designs in clinical trials: why use them, and how to run and report them
Pallmann, P., Bedding, A. W., Choodari-Oskooei, B., Dimairo, M., Flight, L., Hampson, L. V., Holmes, J., Mander, A. P., Odondi, L., Sydes, M. R., et al. (2018) · 2018
Later among the works it cites.
Minimax concave penalized multi-armed bandit model with high-dimensional covariates
Wang, X., Wei, M., and Yao, T. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
18. s997: High dimensional statistics
Rigollet, P. (2015) · 2015
Cited alongside, same era.
Small ball probabilities for linear images of high-dimensional distributions
Rudelson, M. and Vershynin, R. (2015) · 2015
Cited alongside, same era.
Batch learning from logged bandit feedback through counterfactual risk minimization
Swaminathan, A. and Joachims, T. (2015) · 2015
Cited alongside, same era.
Batched bandit problems
Perchet, V., Rigollet, P., Chassang, S., Snowberg, E., et al. (2016) · 2016
Cited alongside, same era.
An information-theoretic analysis of thompson sampling
Russo, D. and Van Roy, B. (2016) · 2016
Cited alongside, same era.
Mostly exploration-free algorithms for contextual bandits
Bastani, H., Bayati, M., and Khosravi, K. (2017) · 2017
Cited alongside, same era.
Behavioral analytics for myopic agents
Mintz, Y., Aswani, A., Kaminsky, P., Flowers, E., and Fukuoka, Y. (2017) · 2017
Cited alongside, same era.
Zhou, M., Fukuoka, Y., Mintz, Y., Goldberg, K., Kaminsky, P., Flowers, E., and Aswani, A. (2018) · 2018
Later among the works it cites.
The big data newsvendor: Practical insights from machine learning
Ban, G.-Y. and Rudin, C. (2019) · 2019
Later among the works it cites.
Batched multi-armed bandits problem
Gao, Z., Han, Y., Ren, Z., and Zhou, Z. (2019) · 2019
Later among the works it cites.
Doubly-robust lasso bandit
Kim, G.-S. and Paik, M. C. (2019) · 2019
Later among the works it cites.
Fast algorithms for online personalized assortment optimization in a big data regime
Miao, S. and Chao, X. (2019) · 2019
Later among the works it cites.
Introduction to multi-armed bandits
Slivkins, A. et al. (2019) · 2019
Later among the works it cites.
High-dimensional statistics: A non-asymptotic viewpoint
Wainwright, M. J. (2019) · 2019
Later among the works it cites.
Online decision making with high-dimensional covariates
Bastani, H. and Bayati, M. (2020) · 2020
Closest in time.
Nonstationary bandits with habituation and recovery dynamics
Mintz, Y., Aswani, A., Kaminsky, P., Flowers, E., and Fukuoka, Y. (2020) · 2020
Closest in time.
Mostly exploration-free algorithms for contextual bandits
Bastani, H., Bayati, M., and Khosravi, K. (2021) · 2021
Closest in time.
Regret lower bound and optimal algorithm for high-dimensional contextual linear bandit
Li, K., Yang, Y., and Narisetty, N. N. (2021) · 2021
Closest in time.
Sparsity-agnostic lasso bandit
Oh, M.-h., Iyengar, G., and Zeevi, A. (2021) · 2021
Closest in time.