Fetching the paper…
Reading the bibliography…
Safe exploration is a challenging and important problem in model-free reinforcement learning (RL).
Constrained Markov decision processes
Eitan Altman · 1999
Earlier work this paper cites.
Nonlinear control synthesis under double constraints
AN Daryin and AB Kurzhanski · 2005
Earlier work this paper cites.
Advanced PID control
Karl Johan Åström · 2006
Earlier work this paper cites.
Feedback systems
Karl Johan Åström and Richard M Murray · 2010
Earlier work this paper cites.
Optimal control
Richard B Vinter · 2010
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Exploration, novelty, surprise, and free energy minimization
Philipp Schwartenbeck, Thomas FitzGerald, Ray Dolan, and Karl Friston · 2013
Earlier work this paper cites.
Reachability-based safe learning with Gaussian processes
Anayo K Akametalu, Jaime F Fisac, Jeremy H Gillula, Shahab Kaynama, Melanie N Zeilinger, and Claire J Tomlin · 2014
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Safe learning of regions of attraction for uncertain, nonlinear systems with Gaussian processes
Felix Berkenkamp, Riccardo Moriconi, Angela P Schoellig, and Andreas Krause · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Safe exploration in finite Markov decision processes with Gaussian processes
Matteo Turchetta, Felix Berkenkamp, and Andreas Krause · 2016
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela Schoellig, and Andreas Krause · 2017
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Earlier work this paper cites.
Uncertainty-aware reinforcement learning for collision avoidance
Gregory Kahn, Adam Villaflor, Vitchyr Pong, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
A Lyapunov-based approach to safe reinforcement learning
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2018
Cited alongside, same era.
Safe exploration in continuous action spaces
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa · 2018
Cited alongside, same era.
Leave no trace: Learning to reset for safe and autonomous reinforcement learning
Benjamin Eysenbach, Shixiang Gu, Julian Ibarz, and Sergey Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
SafePILCO: A software tool for safe and data-efficient policy synthesis
Kyriakos Polymenakos, Nikitas Rontsis, Alessandro Abate, and Stephen Roberts · 2020
Later among the works it cites.
Responsive safety in reinforcement learning by pid lagrangian methods
Adam Stooke, Joshua Achiam, and Pieter Abbeel · 2020
Later among the works it cites.
Safe reinforcement learning via curriculum induction
Matteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause, and Alekh Agarwal · 2020
Later among the works it cites.
Cautious adaptation for reinforcement learning in safety-critical settings
Jesse Zhang, Brian Cheung, Chelsea Finn, Sergey Levine, and Dinesh Jayaraman · 2020
Later among the works it cites.
Conservative safety critics for exploration
Homanga Bharadhwaj, Aviral Kumar, Nicholas Rhinehart, Sergey Levine, Florian Shkurti, and Animesh Garg · 2021
Later among the works it cites.
DESTA: A framework for safe reinforcement learning with markov games of intervention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data-efficient reinforcement learning with probabilistic model predictive control
Sanket Kamthe and Marc Deisenroth · 2018
Cited alongside, same era.
Learning-based model predictive control for safe exploration
Torsten Koller, Felix Berkenkamp, Matteo Turchetta, and Andreas Krause · 2018
Cited alongside, same era.
Two approaches to stochastic optimal control problems with a final-time expectation constraint
Laurent Pfeiffer · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Safe exploration and optimization of constrained MDPs using Gaussian processes
Akifumi Wachi, Yanan Sui, Yisong Yue, and Masahiro Ono · 2018
Cited alongside, same era.
End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks
Richard Cheng, Gábor Orosz, Richard M Murray, and Joel W Burdick · 2019
Cited alongside, same era.
Safely learning to control the constrained linear quadratic regulator
S. Dean, S. Tu, N. Matni, and B. Recht · 2019
Cited alongside, same era.
David Mguni, Joel Jennings, Taher Jafferjee, Aivar Sootla, Yaodong Yang, Changmin Yu, Usman Islam, Ziyan Wang, and Jun Wang · 2021
Later among the works it cites.
Alwayssafe: Reinforcement learning without safety constraint violations during training
Thiago D Simão, Nils Jansen, and Matthijs TJ Spaan · 2021
Later among the works it cites.
WCSAC: Worst-case soft actor critic for safety-constrained reinforcement learning
Qisong Yang, Thiago D Simão, Simon H Tindemans, and Matthijs TJ Spaan · 2021
Later among the works it cites.
https://www.mathworks.com/help/simulink/slref/anti-windup-control-using-a-pid-controller.html
Anti-windup control using a PID controller, 2022 · 2022
Closest in time.
Constrained policy optimization via bayesian world models
Yarden As, Ilnura Usmanova, Sebastian Curi, and Andreas Krause · 2022
Closest in time.
Reinforcement learning with almost sure constraints
Agustin Castellano, Hancheng Min, Enrique Mallada, and Juan Andrés Bazerque · 2022
Closest in time.
SAMBA: Safe model-based & active reinforcement learning
Alexander I Cowen-Rivers, Daniel Palenicek, Vincent Moens, Mohammed Amin Abdullah, Aivar Sootla, Jun Wang, and Haitham Bou-Ammar · 2022
Closest in time.
Magnetic control of tokamak plasmas through deep reinforcement learning
Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de Las Casas, et al · 2022
Closest in time.
Constrained variational policy optimization for safe reinforcement learning
Zuxin Liu, Zhepeng Cen, Vladislav Isenbaev, Wei Liu, Steven Wu, Bo Li, and Ding Zhao · 2022
Closest in time.
Conservative and adaptive penalty for model-based safe reinforcement learning
Yecheng Jason Ma, Andrew Shen, Osbert Bastani, and Jayaraman Dinesh · 2022
Closest in time.
Sauté rl: Almost surely safe reinforcement learning using state augmentation
Aivar Sootla, Alexander I Cowen-Rivers, Taher Jafferjee, Ziyan Wang, David H Mguni, Jun Wang, and Haitham Ammar · 2022
Closest in time.
Muzero’s first step from research into the real world
MuZero Applied Team · 2022
Closest in time.