Fetching the paper…
Reading the bibliography…
We study sequential decision-making with known rewards and unknown constraints, motivated by situations where the constraints represent expensive-to-evaluate human preferences, such as safe and comfortable driving behavior.
The equivalence of two extremum problems
Kiefer, J. and Wolfowitz, J · 1960
Earlier work this paper cites.
The cross-entropy method: a unified approach to combinatorial optimization, Monte-Carlo simulation, and machine learning
Rubinstein, R. Y. and Kroese, D. P · 2004
Earlier work this paper cites.
Optimal design of experiments
Friedrich, P · 2006
Earlier work this paper cites.
Best arm identification in multi-armed bandits
Audibert, J.-Y., Bubeck, S., and Munos, R · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Active learning
Settles, B · 2012
Earlier work this paper cites.
Active reward learning
Daniel, C., Viering, M., Metz, J., Kroemer, O., and Peters, J · 2014
Earlier work this paper cites.
Bayesian optimization with inequality constraints
Gardner, J. R., Kusner, M. J., Xu, Z. E., Weinberger, K. Q., and Cunningham, J. P · 2014
Earlier work this paper cites.
Bayesian optimization with unknown constraints
Gelbart, M. A., Snoek, J., and Adams, R. P · 2014
Earlier work this paper cites.
Best-arm identification in linear bandits
Soare, M., Lazaric, A., and Munos, R · 2014
Earlier work this paper cites.
Sequential resource allocation in linear stochastic bandits
Soare, M · 2015
Cited alongside, same era.
Safe exploration for optimization with Gaussian processes
Sui, Y., Gotovos, A., Burdick, J., and Krause, A · 2015
Cited alongside, same era.
Concrete problems in AI safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Cited alongside, same era.
A general framework for constrained Bayesian optimization using information-based search
Hernández-Lobato, J. M., Gelbart, M. A., Adams, R. P., Hoffman, M. W., and Ghahramani, Z · 2016
Cited alongside, same era.
On the complexity of best-arm identification in multi-armed bandit models
Kaufmann, E., Cappé, O., and Garivier, A · 2016
Cited alongside, same era.
An optimal algorithm for the thresholding bandit problem
Good arm identification via bandit feedback
Kano, H., Honda, J., Sakamaki, K., Matsuura, K., Nakamura, A., and Sugiyama, M · 2019
Later among the works it cites.
Constrained bayesian optimization with max-value entropy search
Perrone, V., Shcherbatyi, I., Jenatton, R., Archambeau, C., and Seeger, M · 2019
Later among the works it cites.
Asking easy questions: A user-friendly approach to active reward learning
Bıyık, E., Palan, M., Landolfi, N. C., Losey, D. P., and Sadigh, D · 2020
Later among the works it cites.
Safe linear stochastic bandits
Khezeli, K. and Bitar, E · 2020
Later among the works it cites.
Bandit algorithms
Lattimore, T. and Szepesvári, C · 2020
Later among the works it cites.
Constrained cross-entropy method for safe reinforcement learning
Wen, M. and Topcu, U · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Locatelli, A., Gutzeit, M., and Carpentier, A · 2016
Cited alongside, same era.
Conservative contextual linear bandits
Kazerouni, A., Ghavamzadeh, M., Abbasi Yadkori, Y., and Van Roy, B · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
Sadigh, D., Dragan, A. D., Sastry, S., and Seshia, S. A · 2017
Cited alongside, same era.
Linear stochastic bandits under safety constraints
Amani, S., Alizadeh, M., and Thrampoulidis, C · 2019
Cited alongside, same era.
Sequential experimental design for transductive linear bandits
Fiez, T., Jain, L., Jamieson, K. G., and Ratliff, L · 2019
Cited alongside, same era.
Safe linear Thompson sampling with side information
Moradipari, A., Amani, S., Alizadeh, M., and Thrampoulidis, C · 2021
Later among the works it cites.
Stochastic bandits with linear constraints
Pacchiano, A., Ghavamzadeh, M., Bartlett, P., and Jiang, H · 2021
Later among the works it cites.
Best arm identification with safety constraints
Wang, Z., Wagenmaker, A., and Jamieson, K · 2022
Closest in time.