Fetching the paper…
Reading the bibliography…
The contextual bandit literature has traditionally focused on algorithms that address the exploration-exploitation tradeoff.
A note on the equivalence of upper confidence bounds and gittins indices for patient agents
Russo, Daniel. 2019 · 1904
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, William R. 1933 · 1933
Earlier work this paper cites.
Bandit processes and dynamic allocation indices
Gittins, John C. 1979 · 1979
Earlier work this paper cites.
A one-armed bandit problem with a concomitant variable
Woodroofe, Michael. 1979 · 1979
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, Tze Leung, Herbert Robbins. 1985 · 1985
Earlier work this paper cites.
Persistent excitation in adaptive systems
Narendra, Kumpati S, Anuradha M Annaswamy. 1987 · 1987
Earlier work this paper cites.
Generalized linear models (Second edition)
McCullagh, P., J. A. Nelder. 1989 · 1989
Earlier work this paper cites.
Theory of Point Estimation
Lehmann, E.L., G. Casella. 1998 · 1998
Earlier work this paper cites.
Strong consistency of maximum quasi-likelihood estimators in generalized linear models with fixed and adaptive designs
Chen, Kani, Inchi Hu, Zhiliang Ying. 1999 · 1999
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, Peter. 2002 · 2002
Earlier work this paper cites.
One-armed bandit problems with covariates
Sarkar, Jyotirmoy. 1991 · 2002
Earlier work this paper cites.
Optimal aggregation of classifiers in statistical learning
Tsybakov, Alexander B, et al. 2004 · 2004
Earlier work this paper cites.
The epoch-greedy algorithm for contextual multi-armed bandits
Langford, John, Tong Zhang. 2007 · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, Varsha, Thomas P Hayes, Sham M Kakade. 2008 · 2008
Earlier work this paper cites.
Estimation of the warfarin dose with clinical and pharmacogenetic data
Consortium, International Warfarin Pharmacogenetics. 2009 · 2009
Earlier work this paper cites.
Woodroofe’s one-armed bandit problem revisited
Goldenshluger, Alexander, Assaf Zeevi. 2009 · 2009
Earlier work this paper cites.
A structured multiarmed bandit problem and the greedy policy
Mersereau, Adam J, Paat Rusmevichientong, John N Tsitsiklis. 2009 · 2009
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Filippi, Sarah, Olivier Cappe, Aurélien Garivier, Csaba Szepesvári. 2010 · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, Lihong, Wei Chu, John Langford, Robert E Schapire. 2010 · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Yasin, Dávid Pál, Csaba Szepesvári. 2011 · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
Chu, Wei, Lihong Li, Lev Reyzin, Robert Schapire. 2011 · 2011
Cited alongside, same era.
The battle trial: personalizing therapy for lung cancer
Kim, Edward S, Roy S Herbst, Ignacio I Wistuba, J Jack Lee, George R Blumenschein, Anne Tsao, David J Stewart, Marshall E Hicks, Jeremy Erasmus, Sanjay Gupta, et al. 2011 · 2011
Cited alongside, same era.
User-friendly tail bounds for matrix martingales
Tropp, Joel A. 2011 · 2011
Cited alongside, same era.
Dynamic pricing under a general parametric choice model
Broder, Josef, Paat Rusmevichientong. 2012 · 2012
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, Sébastien, Nicolò Cesa-Bianchi. 2012 · 2012
Non-stationary bandits with habituation and recovery dynamics
Mintz, Yonatan, Anil Aswani, Philip Kaminsky, Elena Flowers, Yoshimi Fukuoka. 2017 · 2017
Closest in time.
From ads to interventions: Contextual bandits in mobile health
Tewari, Ambuj, Susan A Murphy. 2017 · 2017
Closest in time.
Learning personalized product recommendations with customer disengagement
Bastani, Hamsa, Pavithra Harsha, Georgia Perakis, Divya Singhvi. 2018 · 2018
Closest in time.
Bietti, Alberto, Alekh Agarwal, John Langford. 2018 · 2018
Closest in time.
Bayesian sequential learning for clinical trials of multiple correlated medical interventions
Chick, Stephen E, Noah Gans, Ozge Yapar. 2018 · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, Shipra, Navin Goyal. 2013 · 2013
Cited alongside, same era.
Simultaneously learning and optimizing using controlled variance pricing
den Boer, Arnoud V, Bert Zwart. 2013 · 2013
Cited alongside, same era.
A linear response bandit problem
Goldenshluger, Alexander, Assaf Zeevi. 2013 · 2013
Cited alongside, same era.
Dynamic pricing with an unknown demand model: Asymptotically optimal semi-myopic policies
Keskin, N Bora, Assaf Zeevi. 2014 · 2014
Cited alongside, same era.
Bounded regret for finite-armed structured bandits
Lattimore, Tor, Rémi Munos. 2014 · 2014
Cited alongside, same era.
Learning to optimize via posterior sampling
Russo, Daniel, Benjamin Van Roy. 2014 · 2014
Cited alongside, same era.
Kallus, Nathan, Angela Zhou. 2018 · 2018
Closest in time.
A smoothed analysis of the greedy algorithm for the linear contextual bandit problem
Kannan, Sampath, Jamie H Morgenstern, Aaron Roth, Bo Waggoner, Zhiwei Steven Wu. 2018 · 2018
Closest in time.
On incomplete learning and certainty-equivalence control
Keskin, N Bora, Assaf Zeevi. 2018 · 2018
Closest in time.
Model-reference adaptive control
Nguyen, Nhan T. 2018 · 2018
Closest in time.
Learning to optimize via information-directed sampling
Russo, Daniel, Benjamin Van Roy. 2018 · 2018
Closest in time.
Mnl-bandit: A dynamic learning approach to assortment selection
Agrawal, Shipra, Vashist Avadhanula, Vineet Goyal, Assaf Zeevi. 2019 · 2019
Closest in time.
Meta dynamic pricing: Learning across experiments
Bastani, Hamsa, David Simchi-Levi, Ruihao Zhu. 2019 · 2019
Closest in time.
Dynamic pricing in high-dimensions
Javanmard, Adel, Hamid Nazerzadeh. 2019 · 2019
Closest in time.
High-Dimensional Statistics: A Non-Asymptotic Viewpoint
Wainwright, Martin J. 2019 · 2019
Closest in time.
How do tumor cytogenetics inform cancer treatments? dynamic risk stratification and precision medicine using multi-armed bandits
Zhou, Zhijin, Yingfei Wang, Hamed Mamani, David G Coffey. 2019 · 2019
Closest in time.
Personalized dynamic pricing with machine learning: High dimensional features and heterogeneous elasticity
Ban, Gah-Yi, N Bora Keskin. 2020 · 2020
Closest in time.
Online decision making with high-dimensional covariates
Bastani, Hamsa, Mohsen Bayati. 2020 · 2020
Closest in time.
Provably optimal algorithms for generalized linear contextual bandits
Li, Lihong, Yu Lu, Dengyong Zhou. 2017 · 2080
Closest in time.