Fetching the paper…
Reading the bibliography…
Recent advances in contextual bandit optimization and reinforcement learning have garnered interest in applying these methods to real-world sequential decision making problems.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
Gaussian Processes for Machine Learning
Carl Edward Rasmussen and Christopher K. I. Williams · 2006
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham Kakade · 2008
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Sarah Filippi, Olivier Cappé, Aurélien Garivier, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Oliver Chapelle and Lihong Li · 2011
Cited alongside, same era.
Automatic ad format selection via contextual bandits
Liang Tang, Romer Rosales, Ajit Singh, and Deepak Agarwal · 2013
Cited alongside, same era.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Qui nonero Candela, Denis X. Charles, D. Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
Resourceful contextual bandits
Ashwinkumar Badanidiyuru, John Langford, and Aleksandrs Slivkins · 2014
Cited alongside, same era.
Bandits with concave rewards and convex knapsacks
Shipra Agrawal and Nikhil Devanur · 2014
Cited alongside, same era.
Conservative contextual linear bandits
Abbas Kazerouni, Mohammad Ghavamzadeh, Yasin Abbasi-Yadkori, and Benjamin Van Roy · 2017
Later among the works it cites.
Thompson sampling for the mnl-bandit
Shipra Agrawal, Vashist Avadhanula, Vineet Goyal, and Assaf Zeevi · 2017
Later among the works it cites.
Reinforcement learning with action-derived rewards for chemotherapy and clinical trial dosing regimen selection
Gregory Yauney and Pratik Shah · 2018
Later among the works it cites.
Data center cooling using model-predictive control
Nevena Lazic, Tyler Lu, Craig Boutilier, Moonkyung Ryu, Eehern Wong, Binz Roy, and Greg Imwalle · 2018
Later among the works it cites.
Covariate-adjusted response-adaptive randomization for multi-arm clinical trials using a modified forward looking gittins index rule
Sofía S Villar and William F Rosenberger · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Online decision-making with high-dimensional covariates
Hamsa Bastani and Mohsen Bayati · 2015
Cited alongside, same era.
Deep kernel learning
Andrew Gordon Wilson, Zhiting Hu, Ruslan Salakhutdinov, and Eric P. Xing · 2016
Cited alongside, same era.
Deep bayesian bandits showdown: An emprical comparison of bayesian deep networks for thompson sampling
Carlos Riquelme, Georeg Tucker, and Jasper Snoek · 2018
Later among the works it cites.
Linear stochastic bandits under safety constraints
Sanae Amani, Mahnoosh Alizadeh, and Christos Thrampoulidis · 2019
Closest in time.
Constrained bayesian optimization with noisy experiments
Benjamin Letham, Brian Karrer, Guilherme Ottoni, and Eytan Bakshy · 2019
Closest in time.