Fetching the paper…
Reading the bibliography…
Safe reinforcement learning (RL) that solves constraint-satisfactory policies provides a promising way to the broader safety-critical applications of RL in real-world problems such as robotics.
2004
Earlier work this paper cites.
I. Szita and C. Szepesvari, “Model-based reinforcement learning with nearly tight exploration complexity bounds,” in Proc. Int. Conf. Mach. Learn. , ser. ICML’10. Madison, WI, USA: Omnipress, Jun. 2010, pp. 1031–1038
2010
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “MuJoCo: A physics engine for model-based control,” in Proc. IEEE/RSJ Int. Conf. Intell. Robots Syst. , 2012, pp. 5026–5033
2012
Earlier work this paper cites.
J. F. Fisac, M. Chen, C. J. Tomlin, and S. S. Sastry, “Reach-avoid problems with time-varying dynamics, targets and constraints,” in Proceedings of the 18th International Conference on Hybrid Systems: Computation and Control . Seattle Washington: ACM, Apr. 2015, pp. 11–20
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
F. Berkenkamp, R. Moriconi, A. P. Schoellig, and A. Krause, “Safe learning of regions of attraction for uncertain, nonlinear systems with Gaussian processes,” in IEEE Proc. Conf. Decis. Control. Las Vegas, NV, USA: IEEE, Dec. 2016, pp. 4661–4666
2016
Earlier work this paper cites.
Y. Chow, M. Ghavamzadeh, L. Janson, and M. Pavone, “Risk-Constrained Reinforcement Learning with Percentile Risk Criteria,” JMLR , vol. 18, no. 1, pp. 6070–6120, 2017
2017
Earlier work this paper cites.
M. G. Bellemare, W. Dabney, and R. Munos, “A Distributional Perspective on Reinforcement Learning,” in Proc. Int. Conf. Mach. Learn. PMLR, Jul. 2017, pp. 449–458
2017
Earlier work this paper cites.
J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained Policy Optimization,” in Proc. Int. Conf. Mach. Learn. PMLR, Jul. 2017, pp. 22–31
2017
Earlier work this paper cites.
C. Tessler, D. J. Mankowitz, and S. Mannor, “Reward Constrained Policy Optimization,” in Proc. Int. Conf. Learn. Repr. , Sep. 2018
2018
Earlier work this paper cites.
K. Chua, R. Calandra, R. McAllister, and S. Levine, “Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models,” in Proc. Advances Neural Inf. Process. Syst. , vol. 31. Curran Associates, Inc., 2018
2018
Earlier work this paper cites.
Y. Luo, H. Xu, Y. Li, Y. Tian, T. Darrell, and T. Ma, “Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees,” in Proc. Int. Conf. Learn. Repr. , Sep. 2018
2018
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction , 2nd ed., ser. Adaptive Computation and Machine Learning Series. Cambridge, Massachusetts: The MIT Press, 2018
2018
Earlier work this paper cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor,” in Proc. Int. Conf. Mach. Learn. PMLR, Jul. 2018, pp. 1861–1870
2018
Earlier work this paper cites.
S. Paternain, L. Chamon, M. Calvo-Fullana, and A. Ribeiro, “Constrained Reinforcement Learning Has Zero Duality Gap,” in Proc. Advances Neural Inf. Process. Syst. , vol. 32, 2019
2019
Earlier work this paper cites.
A. Ray, J. Achiam, and D. Amodei, “Benchmarking Safe Exploration in Deep Reinforcement Learning,” Tech. Rep., 2019
2019
Earlier work this paper cites.
M. Janner, J. Fu, M. Zhang, and S. Levine, “When to Trust Your Model: Model-Based Policy Optimization,” in Proc. Advances Neural Inf. Process. Syst. , vol. 32. Curran Associates, Inc., 2019
2019
Earlier work this paper cites.
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson, “Learning Latent Dynamics for Planning from Pixels,” in Proc. Int. Conf. Mach. Learn. PMLR, May 2019, pp. 2555–2565
2019
Cited alongside, same era.
J. F. Fisac, N. F. Lugovoy, V. Rubies-Royo, S. Ghosh, and C. J. Tomlin, “Bridging hamilton-jacobi safety analysis and reinforcement learning,” in Proc. IEEE Int. Conf. Robot. Autom. IEEE, 2019, pp. 8550–8556. [Online]. Available: https://ieeexplore.ieee.org/document/8794107/
2019
Cited alongside, same era.
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, “Dream to Control: Learning Behaviors by Latent Imagination,” in Proc. Int. Conf. Learn. Repr. , Mar. 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
J. Duan, Y. Guan, S. E. Li, Y. Ren, Q. Sun, and B. Cheng, “Distributional Soft Actor-Critic: Off-Policy Reinforcement Learning for Addressing Value Estimation Errors,” IEEE Trans. Neural Netw. Learning Syst. , pp. 1–15, 2021
2021
Later among the works it cites.
Y. Luo and T. Ma, “Learning Barrier Certificates: Towards Safe Reinforcement Learning with Zero Training-time Violations,” in Proc. Advances Neural Inf. Process. Syst. , vol. 34. Curran Associates, Inc., 2021, pp. 25 621–25 632
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Bharadhwaj, A. Kumar, N. Rhinehart, S. Levine, F. Shkurti, and A. Garg, “Conservative Safety Critics for Exploration,” in Proc. Int. Conf. Learn. Repr. , Mar. 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Berkenkamp, “Safe Exploration in Reinforcement Learning: Theory and Applications in Robotics,” Ph.D. dissertation, ETH Zurich, 2019
2020
Cited alongside, same era.
Z. Liu, Q. Liu, L. Tang, K. Jin, H. Wang, M. Liu, and H. Wang, “Visuomotor Reinforcement Learning for Multirobot Cooperative Navigation,” IEEE Trans. Autom. Sci. Eng. , pp. 1–12, 2021
2021
Cited alongside, same era.
E. Altman, Constrained Markov Decision Processes: Stochastic Modeling , 1st ed. Boca Raton: Routledge, Dec. 2021
2021
Cited alongside, same era.
J. Duan, Z. Liu, S. E. Li, Q. Sun, Z. Jia, and B. Cheng, “Adaptive dynamic programming for nonaffine nonlinear optimal control problem with state constraints,” Neurocomputing , Oct. 2021
2021
Cited alongside, same era.
G. Thomas, Y. Luo, and T. Ma, “Safe Reinforcement Learning by Imagining the Near Future,” in Proc. Advances Neural Inf. Process. Syst. , 2021
2021
Cited alongside, same era.
M. A. Zanger, K. Daaboul, and J. Marius Zöllner, “Safe Continuous Control with Constrained Model-Based Policy Optimization,” in Proc. IEEE/RSJ Int. Conf. Intell. Robots Syst. , Sep. 2021, pp. 3512–3519
2021
Cited alongside, same era.
Q. Yang, T. D. Simao, S. H. Tindemans, and M. T. J. Spaan, “WCSAC: Worst-Case Soft Actor Critic for Safety-Constrained Reinforcement Learning,” in Proc. AAAI Conf. Artif. Intell. , vol. 35, May 2021, pp. 10 639–10 646
2021
Cited alongside, same era.
Z. Liu, H. Zhou, B. Chen, S. Zhong, M. Hebert, and D. Zhao, “Constrained Model-based Reinforcement Learning with Robust Cross-Entropy Method,” Mar. 2021
2021
Cited alongside, same era.
B. Thananjeyan, A. Balakrishna, S. Nair, M. Luo, K. Srinivasan, M. Hwang, J. E. Gonzalez, J. Ibarz, C. Finn, and K. Goldberg, “Recovery RL: Safe Reinforcement Learning With Learned Recovery Zones,” IEEE Robot. Autom. Lett. , vol. 6, no. 3, pp. 4915–4922, Jul. 2021
2021
Later among the works it cites.
W. Zhao, T. He, and C. Liu, “Model-free Safe Control for Zero-Violation Reinforcement Learning,” in Proc. Conf. Robot Learn. , Jun. 2021
2021
Later among the works it cites.
M. Brittain and P. Wei, “Scalable Autonomous Separation Assurance With Heterogeneous Multi-Agent Reinforcement Learning,” IEEE Trans. Autom. Sci. Eng. , pp. 1–12, 2022
2022
Closest in time.
Z. Yan, A. R. Kreidieh, E. Vinitsky, A. M. Bayen, and C. Wu, “Unified Automatic Control of Vehicular Systems With Reinforcement Learning,” IEEE Trans. Autom. Sci. Eng. , pp. 1–16, 2022
2022
Closest in time.
Y. Guan, Y. Ren, Q. Sun, S. E. Li, H. Ma, J. Duan, Y. Dai, and B. Cheng, “Integrated Decision and Control: Toward Interpretable and Computationally Efficient Driving Intelligence,” IEEE Trans. Cybern. , pp. 1–15, 2022
2022
Closest in time.
H. Ma, C. Liu, S. E. Li, S. Zheng, and J. Chen, “Joint Synthesis of Safety Certificate and Safe Control Policy Using Constrained Reinforcement Learning,” in Proceedings of The 4th Annual Learning for Dynamics and Control Conference . PMLR, May 2022, pp. 97–109
2022
Closest in time.
D. Yu, H. Ma, S. Li, and J. Chen, “Reachability Constrained Reinforcement Learning,” in Proc. Int. Conf. Mach. Learn. PMLR, Jun. 2022, pp. 25 636–25 655
2022
Closest in time.
Y. As, I. Usmanova, S. Curi, and A. Krause, “Constrained policy optimization via bayesian world models,” in Proc. Int. Conf. Learn. Repr. , 2022. [Online]. Available: https://openreview.net/forum?id=PRZoSmCinhf
2022
Closest in time.
Q. Yang, T. D. Simão, S. H. Tindemans, and M. T. J. Spaan, “Safety-constrained reinforcement learning with a distributional safety critic,” Mach Learn , Jun. 2022
2022
Closest in time.
D. Kim and S. Oh, “TRC: Trust Region Conditional Value at Risk for Safe Reinforcement Learning,” IEEE Robot. Autom. Lett. , vol. 7, no. 2, pp. 2621–2628, Apr. 2022
2022
Closest in time.
C. Dawson, Z. Qin, S. Gao, and C. Fan, “Safe Nonlinear Control Using Robust Neural Lyapunov-Barrier Functions,” in Proc. Conf. Robot Learn. PMLR, Jan. 2022, pp. 1724–1735
2022
Closest in time.
K. Kang, P. Gradu, J. J. Choi, M. Janner, C. Tomlin, and S. Levine, “Lyapunov density models: Constraining distribution shift in learning-based control,” in Proc. Int. Conf. Mach. Learn. , ser. Proceedings of Machine Learning Research, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato, Eds., vol. 162. PMLR, 17–23 Jul 2022, pp. 10 708–10 733. [Online]. Available: https://proceedings.mlr.press/v162/kang22a.html
2022
Closest in time.