Fetching the paper…
Reading the bibliography…
This paper presents the concept of an adaptive safe padding that forces Reinforcement Learning (RL) to synthesise optimal control policies while ensuring safety during the learning process.
The Fundamentals of Learning
Edward L. Thorndike. 1932 · 1932
Earlier work this paper cites.
Maximum likelihood from incomplete data via the EM algorithm
Arthur P Dempster, Nan M Laird, and Donald B Rubin. 1977 · 1977
Earlier work this paper cites.
The temporal logic of programs. In Foundations of Computer Science
Amir Pnueli. 1977 · 1977
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan. 1992 · 1992
Earlier work this paper cites.
Neuro-dynamic Programming
Dimitri P Bertsekas and John N Tsitsiklis. 1996 · 1996
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto. 1998 · 1998
Earlier work this paper cites.
Constrained Markov decision processes
Eitan Altman. 1999 · 1999
Earlier work this paper cites.
Risk-sensitive and minimax control of discrete-time, finite-state Markov decision processes
Stefano P Coraluppi and Steven I Marcus. 1999 · 1999
Earlier work this paper cites.
A note on the existence of optimal policies in total reward dynamic programs with compact action sets
Rolando Cavazos-Cadena, Eugene A. Feinberg, and Raúl Montes-De-Oca. 2000 · 2000
Earlier work this paper cites.
Risk-sensitive reinforcement learning applied to control under constraints
Peter Geibel and Fritz Wysotzki. 2005 · 2005
Earlier work this paper cites.
An application of reinforcement learning to aerobatic helicopter flight. In Advances in Neural Information Processing Systems
Pieter Abbeel, Adam Coates, Morgan Quigley, and Andrew Y Ng. 2007 · 2007
Earlier work this paper cites.
Principles of Model Checking
Christel Baier, Joost-Pieter Katoen, and Kim Guldstrand Larsen. 2008 · 2008
Earlier work this paper cites.
Differential dynamic logic for hybrid systems
André Platzer. 2008 · 2008
Earlier work this paper cites.
Reinforcement learning in finite MDPs: PAC analysis
Alexander L Strehl, Lihong Li, and Michael L Littman. 2009 · 2009
Earlier work this paper cites.
Markov decision processes with state-dependent discount factors and unbounded rewards/costs
Qingda Wei and Xianping Guo. 2011 · 2011
Cited alongside, same era.
Safe exploration of state and action spaces in reinforcement learning
Javier Garcia and Fernando Fernández. 2012 · 2012
Cited alongside, same era.
Safe exploration in Markov decision processes
Teodor Mihai Moldovan and Pieter Abbeel. 2012 · 2012
Cited alongside, same era.
Reinforcement learning with state-dependent discount factor. In ICDL
Naoto Yoshida, Eiji Uchibe, and Kenji Doya. 2013 · 2013
Cited alongside, same era.
Logically-Constrained Neural Fitted Q-Iteration. In Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems
Mohammadhosein Hasanbeig, Alessandro Abate, and Daniel Kroening. 2019b · 2014
Cited alongside, same era.
Optimizing chemical reactions with deep reinforcement learning
Zhenpeng Zhou, Xiaocheng Li, and Richard N. Zare. 2017 · 2017
Later among the works it cites.
Safe reinforcement learning via shielding. In AAAI
Mohammed Alshiekh, Roderick Bloem, Rüdiger Ehlers, Bettina Könighofer, Scott Niekum, and Ufuk Topcu. 2018 · 2018
Later among the works it cites.
A lyapunov-based approach to safe reinforcement learning. In Advances in neural information processing systems
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh. 2018 · 2018
Later among the works it cites.
Model checking
Edmund M Clarke Jr, Orna Grumberg, Daniel Kroening, Doron Peled, and Helmut Veith. 2018 · 2018
Later among the works it cites.
Verifiably Safe Autonomy for Cyber-Physical Systems
Nathan Fulton. 2018 · 2018
Later among the works it cites.
Safe reinforcement learning via formal methods: Toward safe control through proof and learning. In AAAI
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Safe exploration techniques for reinforcement learning–an overview. In International Workshop on Modelling and Simulation for Autonomous Systems
Martin Pecka and Tomas Svoboda. 2014 · 2014
Cited alongside, same era.
Markov decision processes: Discrete stochastic dynamic programming
Martin L. Puterman. 2014 · 2014
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández. 2015 · 2015
Cited alongside, same era.
Human-level Control Through Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Limit-deterministic Büchi automata for linear temporal logic. In CAV
Salomon Sickert, Javier Esparza, Stefan Jaax, and Jan Křetínskỳ. 2016 · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Safe exploration in finite Markov decision processes with Gaussian processes. In Advances in Neural Information Processing Systems
Matteo Turchetta, Felix Berkenkamp, and Andreas Krause. 2016 · 2016
Cited alongside, same era.
Nathan Fulton and André Platzer. 2018 · 2018
Later among the works it cites.
Logically-constrained reinforcement learning
Mohammadhosein Hasanbeig, Alessandro Abate, and Daniel Kroening. 2018 · 2018
Later among the works it cites.
Shielded decision-making in MDPs
Nils Jansen, Bettina Könighofer, Sebastian Junges, and Roderick Bloem. 2018 · 2018
Later among the works it cites.
Constrained cross-entropy method for safe reinforcement learning. In Advances in Neural Information Processing Systems
Min Wen and Ufuk Topcu. 2018 · 2018
Later among the works it cites.
End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks. In Proceedings of the AAAI Conference on Artificial Intelligence
Richard Cheng, Gábor Orosz, Richard M Murray, and Joel W Burdick. 2019 · 2019
Later among the works it cites.
Verifiably Safe Off-Model Reinforcement Learning. In Tools and Algorithms for the Construction and Analysis of Systems (TACAS)
Nathan Fulton and Andre Platzer. 2019 · 2019
Later among the works it cites.
Certified reinforcement learning with logic guidance
Mohammadhosein Hasanbeig, Alessandro Abate, and Daniel Kroening. 2019a · 2019
Later among the works it cites.
Temporal logic guided safe reinforcement learning using control barrier functions
Xiao Li and Calin Belta. 2019 · 2019
Later among the works it cites.
Rethinking the discount factor in reinforcement learning: A decision theoretic approach
Silviu Pitis. 2019 · 2019
Later among the works it cites.