Fetching the paper…
Reading the bibliography…
We study the sequential batch learning problem in linear contextual bandits with finite action sets, where the decision maker is constrained to split incoming individuals into (at most) a fixed number of batches and can only observe outcomes for the individuals within a batch at the batch's end.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Herbert Robbins · 1952
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
Naoki Abe, Alan W Biermann, and Philip M Long · 2003
Earlier work this paper cites.
A learning approach for interactive marketing to a customer segment
Dimitris Bertsimas and Adam J Mersereau · 2007
Earlier work this paper cites.
Crowdsourcing user studies with mechanical turk
Aniket Kittur, Ed H Chi, and Bongwon Suh · 2008
Earlier work this paper cites.
Introduction to Nonparametric Estimation
A. Tsybakov · 2008
Earlier work this paper cites.
Economic analysis of simulation selection problems
Stephen E Chick and Noah Gans · 2009
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Sarah Filippi, Olivier Cappe, Aurélien Garivier, and Csaba Szepesvári · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Online markov decision processes under bandit feedback
Gergely Neu, Andras Antos, András György, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Linearly parameterized bandits
Paat Rusmevichientong and John N Tsitsiklis · 2010
Earlier work this paper cites.
Nonparametric bandits with covariates
Philippe Rigollet and Assaf Zeevi · 2010
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Earlier work this paper cites.
Efficient optimal learning for contextual bandits
Miroslav Dudik, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
The battle trial: personalizing therapy for lung cancer
Edward S Kim, Roy S Herbst, Ignacio I Wistuba, J Jack Lee, George R Blumenschein, Anne Tsao, David J Stewart, Marshall E Hicks, Jeremy Erasmus, Sanjay Gupta, et al · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck, Nicolo Cesa-Bianchi, et al · 2012
Earlier work this paper cites.
The timing of staffing decisions in hospital operating rooms: incorporating workload heterogeneity into the newsvendor problem
Biyu He, Franklin Dexter, Alex Macario, and Stefanos Zenios · 2012
Earlier work this paper cites.
Topics in random matrix theory
Terence Tao · 2012
Cited alongside, same era.
Estimating optimal treatment regimes from a classification perspective
Baqun Zhang, Anastasios A Tsiatis, Marie Davidian, Min Zhang, and Eric Laber · 2012
Cited alongside, same era.
Estimating individualized treatment rules using outcome weighted learning
Yingqi Zhao, Donglin Zeng, A John Rush, and Michael R Kosorok · 2012
Cited alongside, same era.
Further optimal regret bounds for thompson sampling
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
A linear response bandit problem
Alexander Goldenshluger and Assaf Zeevi · 2013
Cited alongside, same era.
Susan Athey and Stefan Wager · 2017
Later among the works it cites.
Scalable generalized linear bandits: Online computation and hashing
Kwang-Sung Jun, Aniruddha Bhargava, Robert Nowak, and Rebecca Willett · 2017
Later among the works it cites.
Provably optimal algorithms for generalized linear contextual bandits
Lihong Li, Yu Lu, and Dengyong Zhou · 2017
Later among the works it cites.
Provably optimal algorithms for generalized linear contextual bandits
Lihong Li, Yu Lu, and Dengyong Zhou · 2017
Later among the works it cites.
Customer acquisition via display advertising using multi-armed bandit experiments
Eric M Schwartz, Eric T Bradlow, and Peter S Fader · 2017
Later among the works it cites.
Customer acquisition via display advertising using multi-armed bandit experiments
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Online learning under delayed feedback
Pooria Joulani, Andras Gyorgy, and Csaba Szepesvári · 2013
Cited alongside, same era.
Modeling delayed feedback in display advertising
Olivier Chapelle · 2014
Cited alongside, same era.
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Doubly robust learning for estimating individualized treatment with censored data
Ying-Qi Zhao, Donglin Zeng, Eric B Laber, Rui Song, Ming Yuan, and Michael Rene Kosorok · 2014
Cited alongside, same era.
Online decision-making with high-dimensional covariates
Hamsa Bastani and Mohsen Bayati · 2015
Cited alongside, same era.
A statistical learning approach to personalization in revenue management
Xi Chen, Zachary Owen, Clark Pixton, and David Simchi-Levi · 2015
Cited alongside, same era.
Eric M Schwartz, Eric T Bradlow, and Peter S Fader · 2017
Later among the works it cites.
Estimation considerations in contextual bandits
Maria Dimakopoulou, Zhengyuan Zhou, Susan Athey, and Guido Imbens · 2018
Later among the works it cites.
Online network revenue management using thompson sampling
Kris Johnson Ferreira, David Simchi-Levi, and He Wang · 2018
Later among the works it cites.
Deep learning with logged bandit feedback
Thorsten Joachims, Adith Swaminathan, and Maarten de Rijke · 2018
Later among the works it cites.
Who should be treated? empirical welfare maximization methods for treatment choice
Toru Kitagawa and Aleksey Tetenov · 2018
Later among the works it cites.
Confounding-robust policy improvement
Nathan Kallus and Angela Zhou · 2018
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2018
Later among the works it cites.
Offline multi-action policy learning: Generalization and optimization
Zhengyuan Zhou, Susan Athey, and Stefan Wager · 2018
Later among the works it cites.
The big data newsvendor: Practical insights from machine learning
Gah-Yi Ban and Cynthia Rudin · 2019
Later among the works it cites.
Online exp3 learning in adversarial bandits with delayed feedback
Ilai Bistritz, Zhengyuan Zhou, Xi Chen, Nicholas Bambos, and Jose Blanchet · 2019
Later among the works it cites.
Batched multi-armed bandits problem
Zijun Gao, Yanjun Han, Zhimei Ren, and Zhengqing Zhou · 2019
Later among the works it cites.
Introduction to multi-armed bandits
Aleksandrs Slivkins et al · 2019
Later among the works it cites.
High-dimensional statistics: A non-asymptotic viewpoint
Martin J Wainwright · 2019
Later among the works it cites.
Learning in generalized linear contextual bandits with stochastic delays
Zhengyuan Zhou, Renyuan Xu, and Jose Blanchet · 2019
Later among the works it cites.