Fetching the paper…
Reading the bibliography…
Safe exploration is crucial for the real-world application of reinforcement learning (RL).
Introduction to reinforcement learning , volume 2
Sutton, R. S., Barto, A. G., et al · 1998
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Altman, E · 1999
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Gers, F. A., Schmidhuber, J., and Cummins, F · 1999
Earlier work this paper cites.
Lyapunov design for safe reinforcement learning
Perkins, T. J. and Barto, A. G · 2002
Earlier work this paper cites.
Risk-sensitive reinforcement learning applied to control under constraints
Geibel, P. and Wysotzki, F · 2005
Earlier work this paper cites.
Efficient exploration and value function generalization in deterministic systems
Wen, Z. and Van Roy, B · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Earlier work this paper cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Chow, Y., Ghavamzadeh, M., Janson, L., and Pavone, M · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
A lyapunov-based approach to safe reinforcement learning
Chow, Y., Nachum, O., Duenez-Guzman, E., and Ghavamzadeh, M · 2018
Cited alongside, same era.
Safe exploration in continuous action spaces
Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y · 2018
Learning to drive in a day
Kendall, A., Hawke, J., Janz, D., Mazur, P., Reda, D., Allen, J.-M., Lam, V.-D., Bewley, A., and Shah, A · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D · 2019
Later among the works it cites.
Benchmarking safe exploration in deep reinforcement learning
Ray, A., Achiam, J., and Amodei, D · 2019
Later among the works it cites.
Open-sourced reinforcement learning environments for surgical robotics
Richter, F., Orosco, R. K., and Yip, M. C · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
Wang, T., Bao, X., Clavera, I., Hoang, J., Wen, Y., Langlois, E., Zhang, S., Zhang, G., Abbeel, P., and Ba, J · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Van Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
On first-order meta-learning algorithms
Nichol, A., Achiam, J., and Schulman, J · 2018
Cited alongside, same era.
Time limits in reinforcement learning
Pardo, F., Tavakoli, A., Levdik, V., and Kormushev, P · 2018
Cited alongside, same era.
Constrained exploration and recovery from experience shaping
Pham, T.-H., De Magistris, G., Agravante, D. J., Chaudhury, S., Munawar, A., and Tachibana, R · 2018
Cited alongside, same era.
Trial without error: Towards safe reinforcement learning via human intervention
Saunders, W., Sastry, G., Stuhlmueller, A., and Evans, O · 2018
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2019
Cited alongside, same era.
Bharadhwaj, H., Kumar, A., Rhinehart, N., Levine, S., Shkurti, F., and Garg, A · 2020
Later among the works it cites.
Ipo: Interior-point policy optimization under constraints
Liu, Y., Ding, J., and Liu, X · 2020
Later among the works it cites.
Lyapunov barrier policy optimization
Sikchi, H., Zhou, W., and Held, D · 2020
Later among the works it cites.
Learning to be safe: Deep rl with a safety critic
Srinivasan, K., Eysenbach, B., Ha, S., Tan, J., and Finn, C · 2020
Later among the works it cites.
Responsive safety in reinforcement learning by pid lagrangian methods
Stooke, A., Achiam, J., and Abbeel, P · 2020
Later among the works it cites.
Learning for safety-critical control with control barrier functions
Taylor, A., Singletary, A., Yue, Y., and Ames, A · 2020
Later among the works it cites.
Online constrained model-based reinforcement learning
Van Niekerk, B., Damianou, A., and Rosman, B · 2020
Later among the works it cites.
Safe reinforcement learning in constrained markov decision processes
Wachi, A. and Sui, Y · 2020
Later among the works it cites.
Cautious adaptation for reinforcement learning in safety-critical settings
Zhang, J., Cheung, B., Finn, C., Levine, S., and Jayaraman, D · 2020
Later among the works it cites.