Fetching the paper…
Reading the bibliography…
In this paper, we study a family of conservative bandit problems (CBPs) with sample-path reward constraints, i.e., the learner's reward performance must be at least as well as a given baseline at any time.
Contextual Combinatorial Conservative Bandits
Zhang, X.; Li, S.; and Liu, W. 2019 · 1911
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. 1933 · 1933
Earlier work this paper cites.
Portfolio Selection
Markowitz, H. M.; et al. 1952 · 1952
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P.; Cesa-Bianchi, N.; and Fischer, P. 2002 · 2002
Earlier work this paper cites.
Stochastic Linear Optimization under Bandit Feedback
Dani, V.; Hayes, T.; and Kakade, S. M. 2008 · 2008
Earlier work this paper cites.
Improved Algorithms for Linear Stochastic Bandits
Abbasi-yadkori, Y.; Pál, D.; and Szepesvári, C. 2011 · 2011
Earlier work this paper cites.
Analysis of thompson sampling for the multi-armed bandit problem
Agrawal, S.; and Goyal, N. 2012 · 2012
Earlier work this paper cites.
A One-Size-Fits-All Solution to Conservative Bandit Problems
Du, Y.; Wang, S.; and Huang, L. 2020 · 2012
Cited alongside, same era.
Risk-aversion in multi-armed bandits
Sani, A.; Lazaric, A.; and Munos, R. 2012 · 2012
Cited alongside, same era.
Bounded regret in stochastic multi-armed bandits
Bubeck, S.; Perchet, V.; and Rigollet, P. 2013 · 2013
Cited alongside, same era.
Robust risk-averse stochastic multi-armed bandits
Maillard, O.-A. 2013 · 2013
Cited alongside, same era.
Contextual Combinatorial Bandit and its Application on Diversified Online Recommendation
Qin, L.; Chen, S.; and Zhu, X. 2014 · 2014
Cited alongside, same era.
An optimal algorithm for the Thresholding Bandit Problem
Locatelli, A.; Gutzeit, M.; and Carpentier, A. 2016 · 2016
Conservative contextual linear bandits
Kazerouni, A.; Ghavamzadeh, M.; Yadkori, Y. A.; and Van Roy, B. 2017 · 2017
Later among the works it cites.
Linear Stochastic Bandits Under Safety Constraints
Amani, S.; Alizadeh, M.; and Thrampoulidis, C. 2019 · 2019
Later among the works it cites.
Risk-averse stochastic convex bandit
Cardoso, A. R.; and Xu, H. 2019 · 2019
Later among the works it cites.
Conservative Exploration using Interleaving
Katariya, S.; Kveton, B.; Wen, Z.; and Potluru, V. 2019 · 2019
Later among the works it cites.
Decision variance in risk-averse online learning
Vakili, S.; Boukouvalas, A.; and Zhao, Q. 2019 · 2019
Later among the works it cites.
Improved Algorithms for Conservative Exploration in Bandits
Garcelon, E.; Ghavamzadeh, M.; Lazaric, A.; and Pirotta, M. 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Conservative bandits
Wu, Y.; Shariff, R.; Lattimore, T.; and Szepesvári, C. 2016 · 2016
Cited alongside, same era.
Safe Linear Stochastic Bandits
Khezeli, K.; and Bitar, E. 2020 · 2020
Closest in time.