Fetching the paper…
Reading the bibliography…
The design and performance analysis of bandit algorithms in the presence of stage-wise safety or reliability constraints has recently garnered significant interest.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P · 2002
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V., Hayes, T. P., and Kakade, S. M · 2008
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E · 2010
Earlier work this paper cites.
Linearly parameterized bandits
Rusmevichientong, P. and Tsitsiklis, J. N · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Chu, W., Li, L., Reyzin, L., and Schapire, R · 2011
Earlier work this paper cites.
Analysis of thompson sampling for the multi-armed bandit problem
Agrawal, S. and Goyal, N · 2012
Earlier work this paper cites.
Thompson sampling: An asymptotically optimal finite-time analysis
Kaufmann, E., Korda, N., and Munos, R · 2012
Earlier work this paper cites.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, S. and Goyal, N · 2013
Earlier work this paper cites.
Provably safe and robust learning-based model predictive control
Aswani, A., Gonzalez, H., Sastry, S. S., and Tomlin, C · 2013
Cited alongside, same era.
Concentration inequalities: A nonasymptotic theory of independence
Boucheron, S., Lugosi, G., and Massart, P · 2013
Cited alongside, same era.
Reachability-based safe learning with gaussian processes
Akametalu, A. K., Fisac, J. F., Gillula, J. H., Kaynama, S., Zeilinger, M. N., and Tomlin, C. J · 2014
Cited alongside, same era.
Thompson sampling for complex online problems
Gopalan, A., Mannor, S., and Mansour, Y · 2014
Cited alongside, same era.
Learning to optimize via posterior sampling
Russo, D. and Van Roy, B · 2014
Cited alongside, same era.
Thompson sampling for learning parameterized markov decision processes
Gopalan, A. and Mannor, S · 2015
Cited alongside, same era.
An information-theoretic analysis of thompson sampling
Russo, D. and Van Roy, B · 2016
Later among the works it cites.
Linear thompson sampling revisited
Abeille, M., Lazaric, A., et al · 2017
Later among the works it cites.
Conservative contextual linear bandits
Kazerouni, A., Ghavamzadeh, M., Abbasi, Y., and Van Roy, B · 2017
Later among the works it cites.
An information-theoretic analysis for thompson sampling with many actions
Dong, S. and Van Roy, B · 2018
Later among the works it cites.
Learning-based model predictive control for safe exploration
Koller, T., Berkenkamp, F., Turchetta, M., and Krause, A · 2018
Later among the works it cites.
Stagewise safe bayesian optimization with gaussian processes
Sui, Y., Burdick, J., Yue, Y., et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bootstrapped thompson sampling and deep exploration
Osband, I. and Van Roy, B · 2015
Cited alongside, same era.
Safe exploration for optimization with gaussian processes
Sui, Y., Gotovos, A., Burdick, J. W., and Krause, A · 2015
Cited alongside, same era.
Robust constrained learning-based nmpc enabling reliable mobile robot path tracking
Ostafew, C. J., Schoellig, A. P., and Barfoot, T. D · 2016
Cited alongside, same era.
Amani, S., Alizadeh, M., and Thrampoulidis, C · 2019
Closest in time.
On the performance of thompson sampling on logistic bandits
Dong, S., Ma, T., and Van Roy, B · 2019
Closest in time.
Safe convex learning under uncertain constraints
Usmanova, I., Krause, A., and Kamgarpour, M · 2019
Closest in time.