Fetching the paper…
Reading the bibliography…
In this paper we propose a formal, model-agnostic meta-learning framework for safe reinforcement learning.
Temporal logic guided safe reinforcement learning using control barrier functions
Li, X.; and Belta, C. 2019 · 1903
Earlier work this paper cites.
System development , volume 85
Jackson, M. A.; and Cameron, J. 1983 · 1983
Earlier work this paper cites.
Seeing the child, knowing the person
Balaban, N. 1995 · 1995
Earlier work this paper cites.
Reinforcement Learning: An Introduction , volume 1
Sutton, R. S.; and Barto, A. G. 1998 · 1998
Earlier work this paper cites.
Constrained Markov decision processes: stochastic modeling
Altman, E. 1999 · 1999
Earlier work this paper cites.
Risk-sensitive and minimax control of discrete-time, finite-state Markov decision processes
Coraluppi, S. P.; and Marcus, S. I. 1999 · 1999
Earlier work this paper cites.
Essentials of stochastic processes , volume 1
Durrett, R. 1999 · 1999
Earlier work this paper cites.
Computational techniques for the verification of hybrid systems
Tomlin, C. J.; Mitchell, I.; Bayen, A. M.; and Oishi, M. 2003 · 2003
Earlier work this paper cites.
Risk-sensitive reinforcement learning applied to control under constraints
Geibel, P.; and Wysotzki, F. 2005 · 2005
Earlier work this paper cites.
An application of reinforcement learning to aerobatic helicopter flight
Abbeel, P.; Coates, A.; Quigley, M.; and Ng, A. 2006 · 2006
Earlier work this paper cites.
Barrier certificates for nonlinear model validation
Prajna, S. 2006 · 2006
Earlier work this paper cites.
A Composable Specification Language for Reinforcement Learning Tasks
Jothimurugan, K.; Alur, R.; and Bastani, O. 2020 · 2008
Earlier work this paper cites.
Interaction between learning and development
Vygotsky, L.; et al. 2011 · 2011
Earlier work this paper cites.
Safe exploration of state and action spaces in reinforcement learning
Garcia, J.; and Fernández, F. 2012 · 2012
Earlier work this paper cites.
Safe exploration in Markov decision processes
Moldovan, T. M.; and Abbeel, P. 2012 · 2012
Earlier work this paper cites.
Policy gradients with variance related risk criteria
Tamar, A.; Di Castro, D.; and Mannor, S. 2012 · 2012
Earlier work this paper cites.
Scaling up robust MDPs by reinforcement learning
Tamar, A.; Xu, H.; and Mannor, S. 2013 · 2013
Earlier work this paper cites.
Logically-Constrained Neural Fitted Q-Iteration
Hasanbeig, H.; Abate, A.; and Kroening, D. 2019 · 2014
Earlier work this paper cites.
Safe exploration techniques for reinforcement learning–an overview
Pecka, M.; and Svoboda, T. 2014 · 2014
Cited alongside, same era.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L. 2014 · 2014
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Garcıa, J.; and Fernández, F. 2015 · 2015
Cited alongside, same era.
Concrete problems in AI safety
Amodei, D.; Olah, C.; Steinhardt, J.; Christiano, P.; Schulman, J.; and Mané, D. 2016 · 2016
Cited alongside, same era.
Blundell, C.; Uria, B.; Pritzel, A.; Li, Y.; Ruderman, A.; Leibo, J. Z.; Rae, J.; Wierstra, D.; and Hassabis, D. 2016 · 2016
Cited alongside, same era.
Batch policy learning under constraints
Le, H.; Voloshin, C.; and Yue, Y. 2019 · 2019
Later among the works it cites.
Reinforcement learning with convex constraints
Miryoosefi, S.; Brantley, K.; Daume III, H.; Dudik, M.; and Schapire, R. E. 2019 · 2019
Later among the works it cites.
Neurosymbolic reinforcement learning with formally verified exploration
Anderson, G.; Verma, A.; Dillig, I.; and Chaudhuri, S. 2020 · 2020
Later among the works it cites.
Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement Learning
Petrenko, A.; Huang, Z.; Kumar, T.; Sukhatme, G.; and Koltun, V. 2020 · 2020
Later among the works it cites.
Conservative Safety Critics for Exploration
Bharadhwaj, H.; Kumar, A.; Rhinehart, N.; Levine, S.; Shkurti, F.; and Garg, A. 2021 · 2021
Later among the works it cites.
DeepSynth: Program Synthesis for Automatic Task Segmentation in Deep Reinforcement Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lipton, Z. C.; Azizzadenesheli, K.; Kumar, A.; Li, L.; Gao, J.; and Deng, L. 2016 · 2016
Cited alongside, same era.
Stabilization with guaranteed safety using control Lyapunov–barrier function
Romdlony, M. Z.; and Jayawardhana, B. 2016 · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D.; Huang, A.; Maddison, C. J.; Guez, A.; Sifre, L.; van den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; et al. 2016 · 2016
Cited alongside, same era.
Safe exploration in finite Markov decision processes with Gaussian processes
Turchetta, M.; Berkenkamp, F.; and Krause, A. 2016 · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J.; Held, D.; Tamar, A.; and Abbeel, P. 2017 · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C.; Abbeel, P.; and Levine, S. 2017 · 2017
Cited alongside, same era.
Neural episodic control
Pritzel, A.; Uria, B.; Srinivasan, S.; Badia, A. P.; Vinyals, O.; Hassabis, D.; Wierstra, D.; and Blundell, C. 2017 · 2017
Cited alongside, same era.
Hasanbeig, H.; Yogananda Jeppu, N.; Abate, A.; Melham, T.; and Kroening, D. 2021 · 2021
Later among the works it cites.
Verifiably safe exploration for end-to-end reinforcement learning
Hunt, N.; Fulton, N.; Magliacane, S.; Hoang, T. N.; Das, S.; and Solar-Lezama, A. 2021 · 2021
Later among the works it cites.
Safe Reinforcement Learning via Shielding under Partial Observability
Carr, S.; Jansen, N.; Junges, S.; and Topcu, U. 2022 · 2022
Later among the works it cites.
LCRL: Certified Policy Synthesis via Logically-Constrained Reinforcement Learning
Hasanbeig, H.; Kroening, D.; and Abate, A. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022 · 2022
Later among the works it cites.
Safe policies for reinforcement learning via primal-dual methods
Paternain, S.; Calvo-Fullana, M.; Chamon, L. F.; and Ribeiro, A. 2022 · 2022
Later among the works it cites.
ViZDoom Competitions: Playing Doom from Pixels
Wydmuch, M.; Kempka, M.; and Jaśkowski, W. 2019 · 2022
Later among the works it cites.
Temporal Logic Guided Meta Q-Learning of Multiple Tasks
Zhang, H.; and Kan, Z. 2022 · 2022
Later among the works it cites.
Certified Reinforcement Learning with Logic Guidance
Hasanbeig, H.; Kroening, D.; and Abate, A. 2023 · 2023
Later among the works it cites.
Wang, J.; Hasanbeig, H.; Tan, K.; Sun, Z.; and Kantaros, Y. 2023 · 2023
Later among the works it cites.
Symbolic Task Inference in Deep Reinforcement Learning
Hasanbeig, H.; Yogananda Jeppu, N.; Abate, A.; Melham, T.; and Kroening, D. 2024 · 2024
Closest in time.
Safeguarded Progress in Reinforcement Learning: Safe Bayesian Exploration for Control Policy Synthesis
Mitta, R.; Hasanbeig, H.; Wang, J.; Kroening, D.; Kantaros, Y.; and Abate, A. 2024 · 2024
Closest in time.