Fetching the paper…
Reading the bibliography…
We study a constrained contextual linear bandit setting, where the goal of the agent is to produce a sequence of policies, whose expected cumulative reward over the course of $T$ rounds is maximum, and each has an expected cost below a certain threshold $\tau$.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. Thompson · 1933
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. Lai and H. Robbins · 1985
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
V. Dani, T. Hayes, and S. Kakade · 2008
Earlier work this paper cites.
Application of multi-armed bandits to sensor management
R. Washburn · 2008
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. Schapire · 2010
Earlier work this paper cites.
Linearly parameterized bandits
P. Rusmevichientong and J. Tsitsiklis · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
Bandits with knapsacks
A. Badanidiyuru, R. Kleinberg, and A. Slivkins · 2013
Earlier work this paper cites.
The combinatorial multi-armed bandit problem and its application to real-time strategy games
S. Ontanón · 2013
Cited alongside, same era.
Bandits with concave rewards and convex knapsacks
S. Agrawal and N. Devanur · 2014
Cited alongside, same era.
Resourceful contextual bandits
A. Badanidiyuru, J. Langford, and A. Slivkins · 2014
Cited alongside, same era.
Multi-armed bandit models for the optimal design of clinical trials: Benefits and challenges
S. Villar, J. Bowden, and J. Wason · 2015
Cited alongside, same era.
Algorithms with logarithmic or sub-linear regret for constrained contextual bandits
H. Wu, R. Srikant, X. Liu, and C. Jiang · 2015
Cited alongside, same era.
Linear contextual bandits with knapsacks
S. Agrawal and N. Devanur · 2016
Cited alongside, same era.
Conservative bandits
Y. Wu, R. Shariff, T. Lattimore, and C. Szepesvári · 2016
Later among the works it cites.
Conservative contextual linear bandits
A. Kazerouni, M. Ghavamzadeh, Y. Abbasi Yadkori, and B. Van Roy · 2017
Later among the works it cites.
Using contextual bandits with behavioral constraints for constrained online movie recommendation
A. Balakrishnan, D. Bouneffouf, N. Mattei, and F. Rossi · 2018
Later among the works it cites.
A tutorial on Thompson sampling
D. Russo, B. Van Roy, A. Kazerouni, I. Osband, and Z. Wen · 2018
Later among the works it cites.
Linear stochastic bandits under safety constraints
S. Amani, M. Alizadeh, and C. Thrampoulidis · 2019
Later among the works it cites.
Bandit Algorithms
T. Lattimore and C. Szepesvári · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the complexity of best-arm identification in multi-armed bandit models
Emilie Kaufmann, Olivier Cappé, and Aurélien Garivier · 2016
Cited alongside, same era.
Multi-armed bandits with application to 5G small cells
S. Maghsudi and E. Hossain · 2016
Cited alongside, same era.
Further optimal regret bounds for Thompson sampling
S. Agrawal and N. Goyal
Cited in the paper.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawal and N. Goyal
Cited in the paper.
A. Moradipari, S. Amani, M. Alizadeh, and C. Thrampoulidis · 2019
Later among the works it cites.
Improved algorithms for conservative exploration in bandits
E. Garcelon, M. Ghavamzadeh, A. Lazaric, and M. Pirotta · 2020
Closest in time.