Fetching the paper…
Reading the bibliography…
Satisfying safety constraints almost surely (or with probability one) can be critical for the deployment of Reinforcement Learning (RL) in real-life applications.
Discrete-time Markov control processes with discounted unbounded costs: optimality criteria
Hernández-Lerma, O. and Muñoz de Ozak, M · 1992
Earlier work this paper cites.
Safe active learning for time-series modeling with Gaussian processes
Zimmer, C., Meister, M., and Nguyen-Tuong, D · 1992
Earlier work this paper cites.
Discrete-time controlled Markov processes with average cost criterion: a survey
Arapostathis, A., Borkar, V. S., Fernández-Gaucherand, E., Ghosh, M. K., and Marcus, S. I · 1993
Earlier work this paper cites.
Nonlinear programming
Bertsekas, D. P · 1997
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Altman, E · 1999
Earlier work this paper cites.
Optimization of conditional value-at-risk
Rockafellar, R. T., Uryasev, S., et al · 2000
Earlier work this paper cites.
Nonlinear control synthesis under double constraints
Daryin, A. and Kurzhanski, A · 2005
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
Rasmussen, C. E. and Williams, C. K. I · 2005
Earlier work this paper cites.
A Markov decision model for a surveillance application and risk-sensitive Markov decision processes, 2010
Ott, J. T · 2010
Earlier work this paper cites.
Markov decision processes with average-value-at-risk criteria
Bäuerle, N. and Ott, J · 2011
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Earlier work this paper cites.
Discrete-time Markov control processes: basic optimality criteria , volume 30
Hernández-Lerma, O. and Lasserre, J. B · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Reachability-based safe learning with Gaussian processes
Akametalu, A. K., Fisac, J. F., Gillula, J. H., Kaynama, S., Zeilinger, M. N., and Tomlin, C. J · 2014
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Optimizing the cvar via sampling
Tamar, A., Glassner, Y., and Mannor, S · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Safe exploration in finite Markov decision processes with Gaussian processes
Turchetta, M., Berkenkamp, F., and Krause, A · 2016
Earlier work this paper cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Earlier work this paper cites.
Safe model-based reinforcement learning with stability guarantees
Berkenkamp, F., Turchetta, M., Schoellig, A., and Krause, A · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Chow, Y., Ghavamzadeh, M., Janson, L., and Pavone, M · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
A Lyapunov-based approach to safe reinforcement learning
Chow, Y., Nachum, O., Duenez-Guzman, E., and Ghavamzadeh, M · 2018
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Later among the works it cites.
Barrier-certified adaptive reinforcement learning with applications to brushbot navigation
Ohnishi, M., Wang, L., Notomista, G., and Egerstedt, M · 2019
Later among the works it cites.
Benchmarking safe exploration in deep reinforcement learning, 2019
Ray, A., Achiam, J., and Amodei, D · 2019
Later among the works it cites.
Projection-based constrained policy optimization
Yang, T.-Y., Rosca, J., Narasimhan, K., and Ramadge, P. J · 2019
Later among the works it cites.
Conservative safety critics for exploration
Bharadhwaj, H., Kumar, A., Rhinehart, N., Levine, S., Shkurti, F., and Garg, A · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y · 2018
Cited alongside, same era.
Leave no trace: Learning to reset for safe and autonomous reinforcement learning
Eysenbach, B., Gu, S., Ibarz, J., and Levine, S · 2018
Cited alongside, same era.
A general safety framework for learning-based control in uncertain robotic systems
Fisac, J. F., Akametalu, A. K., Zeilinger, M. N., Kaynama, S., Gillula, J., and Tomlin, C. J · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Data-efficient reinforcement learning with probabilistic model predictive control
Kamthe, S. and Deisenroth, M · 2018
Cited alongside, same era.
Learning-based model predictive control for safe exploration
Koller, T., Berkenkamp, F., Turchetta, M., and Krause, A · 2018
Cited alongside, same era.
Simple random search provides a competitive approach to reinforcement learning
Mania, H., Guy, A., and Recht, B · 2018
Cited alongside, same era.
Ding, D., Zhang, K., Basar, T., and Jovanovic, M. R · 2020
Later among the works it cites.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Lee, A., Nagabandi, A., Abbeel, P., and Levine, S · 2020
Later among the works it cites.
SafePILCO: A software tool for safe and data-efficient policy synthesis
Polymenakos, K., Rontsis, N., Abate, A., and Roberts, S · 2020
Later among the works it cites.
Responsive safety in reinforcement learning by pid lagrangian methods
Stooke, A., Achiam, J., and Abbeel, P · 2020
Later among the works it cites.
Safe reinforcement learning via curriculum induction
Turchetta, M., Kolobov, A., Shah, S., Krause, A., and Agarwal, A · 2020
Later among the works it cites.
Calvo-Fullana, M., Paternain, S., Chamon, L. F., and Ribeiro, A · 2021
Later among the works it cites.
DESTA: A framework for safe reinforcement learning with markov games of intervention
Mguni, D., Jennings, J., Jafferjee, T., Sootla, A., Yang, Y., Yu, C., Islam, U., Wang, Z., and Wang, J · 2021
Later among the works it cites.
Mbrl-lib: A modular library for model-based reinforcement learning
Pineda, L., Amos, B., Zhang, A., Lambert, N. O., and Calandra, R · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., and Dormann, N · 2021
Later among the works it cites.
Seaborn: statistical data visualization
Waskom, M. L · 2021
Later among the works it cites.
WCSAC: Worst-case soft actor critic for safety-constrained reinforcement learning
Yang, Q., Simão, T. D., Tindemans, S. H., and Spaan, M. T · 2021
Later among the works it cites.
SAMBA: Safe model-based & active reinforcement learning
Cowen-Rivers, A. I., Palenicek, D., Moens, V., Abdullah, M. A., Sootla, A., Wang, J., and Bou-Ammar, H · 2022
Closest in time.
Sauté RL: Almost surely safe reinforcement learning using state augmentation, 2022
Sootla, A., Cowen-Rivers, A. I., Jafferjee, T., and Wang, Z · 2022
Closest in time.