Fetching the paper…
Reading the bibliography…
Safe reinforcement learning has been a promising approach for optimizing the policy of an agent that operates in safety-critical applications.
Constrained model predictive control: Stability and optimality
Mayne, D. Q., Rawlings, J. B., Rao, C. V., and Scokaert, P. O · 2000
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Schölkopf, B. and Smola, A. J · 2001
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
Gaussian processes in machine learning
Rasmussen, C. E · 2004
Earlier work this paper cites.
A theoretical analysis of model-based interval estimation
Strehl, A. L. and Littman, M. L · 2005
Earlier work this paper cites.
PAC model-free reinforcement learning
Strehl, A. L., Li, L., Wiewiora, E., Langford, J., and Littman, M. L · 2006
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Auer, P. and Ortner, R · 2007
Earlier work this paper cites.
Mars reconnaissance orbiter’s high resolution imaging science experiment (HiRISE)
McEwen, A. S., Eliason, E. M., Bergstrom, J. W., et al · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for Markov decision processes
Strehl, A. L. and Littman, M. L · 2008
Cited alongside, same era.
Near-Bayesian exploration in polynomial time
Kolter, J. Z. and Ng, A. Y · 2009
Cited alongside, same era.
Gaussian process optimization in the bandit setting: No regret and experimental design
Srinivas, N., Krause, A., Kakade, S. M., and Seeger, M · 2010
Cited alongside, same era.
Near-optimal BRL using optimistic local transitions
Araya, M., Buffet, O., and Thomas, V · 2012
Cited alongside, same era.
Provably safe and robust learning-based model predictive control
Aswani, A., Gonzalez, H., Sastry, S. S., and Tomlin, C · 2013
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F · 2015
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Later among the works it cites.
Safe model-based reinforcement learning with stability guarantees
Berkenkamp, F., Turchetta, M., Schoellig, A., and Krause, A · 2017
Later among the works it cites.
On kernelized multi-armed bandits
Chowdhury, S. R. and Gopalan, A · 2017
Later among the works it cites.
A general safety framework for learning-based control in uncertain robotic systems
Fisac, J. F., Akametalu, A. K., Zeilinger, M. N., Kaynama, S., Gillula, J., and Tomlin, C. J · 2018
Later among the works it cites.
Stagewise safe Bayesian optimization with Gaussian processes
Sui, Y., Zhuang, V., Burdick, J. W., and Yue, Y · 2018
Later among the works it cites.
Lyapunov-based safe policy optimization for continuous control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Safe exploration for optimization with Gaussian processes
Sui, Y., Gotovos, A., Burdick, J. W., and Krause, A · 2015
Cited alongside, same era.
Safe exploration in finite Markov decision processes with Gaussian processes
Turchetta, M., Berkenkamp, F., and Krause, A · 2016
Cited alongside, same era.
Safe exploration and optimization of constrained MDPs using Gaussian processes
Wachi, A., Sui, Y., Yue, Y., and Ono, M · 2016
Cited alongside, same era.
Chow, Y., Nachum, O., Faust, A., Duenez-Guzman, E., and Ghavamzadeh, M · 2019
Later among the works it cites.
Benchmarking safe exploration in deep reinforcement learning
Ray, A., Achiam, J., and Amodei, D · 2019
Later among the works it cites.
Safe exploration for interactive machine learning
Turchetta, M., Berkenkamp, F., and Krause, A · 2019
Later among the works it cites.