Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) involves performing exploratory actions in an unknown system.
A note on dijkstra’s shortest path algorithm
Donald B Johnson · 1973
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Tommi Jaakkola, Michael I Jordan, and Satinder P Singh · 1994
Earlier work this paper cites.
Physiology of diving of birds and mammals
Patrick J Butler and David R Jones · 1997
Earlier work this paper cites.
Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program
Eitan Altman · 1998
Earlier work this paper cites.
Constrained Markov Decision Processes
Eitan Altman · 1999
Earlier work this paper cites.
Optimal stopping of markov processes: Hilbert space theory, approximation algorithms, and an application to pricing high-dimensional financial derivatives
John N Tsitsiklis and Benjamin Van Roy · 1999
Earlier work this paper cites.
Safe exploration for reinforcement learning
Alexander Hans, Daniel Schneegaß, Anton Maximilian Schäfer, and Steffen Udluft · 2008
Earlier work this paper cites.
Learning to control a low-cost manipulator using data-efficient reinforcement learning
Marc Peter Deisenroth, Carl Edward Rasmussen, and Dieter Fox · 2011
Earlier work this paper cites.
Adaptive algorithms and stochastic approximations
Albert Benveniste, Michel Métivier, and Pierre Priouret · 2012
Earlier work this paper cites.
Approximate dynamic programming
Dimitri P Bertsekas · 2012
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier García, Fern, and o Fernández · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Convex synthesis of randomized policies for controlled markov chains with density safety upper bound constraints
Mahmoud El Chamie, Yue Yu, and Behçet Açıkmeşe · 2016
Earlier work this paper cites.
Constrained policy optimization, 2017
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela P. Schoellig, and Andreas Krause · 2017
Earlier work this paper cites.
Peng Peng, Ying Wen, Yaodong Yang, Quan Yuan, Zhenkun Tang, Haitao Long, and Jun Wang · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
A lyapunov-based approach to safe reinforcement learning
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2018
Cited alongside, same era.
How many random seeds? statistical power analysis in deep reinforcement learning experiments
C. Colas, O. Sigaud, and P. Y. Oudeyer · 2018
Cited alongside, same era.
Samba: Safe model-based & active reinforcement learning, 2020
Alexander I. Cowen-Rivers, Daniel Palenicek, Vincent Moens, Mohammed Abdullah, Aivar Sootla, Jun Wang, and Haitham Ammar · 2020
Later among the works it cites.
Responsive safety in reinforcement learning by pid lagrangian methods
Adam Stooke, Joshua Achiam, and Pieter Abbeel · 2020
Later among the works it cites.
Safe reinforcement learning via curriculum induction
Matteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause, and Alekh Agarwal · 2020
Later among the works it cites.
An overview of multi-agent reinforcement learning from game theoretical perspective
Yaodong Yang and Jun Wang · 2020
Later among the works it cites.
Conservative safety critics for exploration
Homanga Bharadhwaj, Aviral Kumar, Nicholas Rhinehart, Sergey Levine, Florian Shkurti, and Animesh Garg · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa · 2018
Cited alongside, same era.
Leave no trace: Learning to reset for safe and autonomous reinforcement learning
Benjamin Eysenbach, Shixiang Gu, Julian Ibarz, and Sergey Levine · 2018
Cited alongside, same era.
Learning-based model predictive control for safe exploration
Torsten Koller, Felix Berkenkamp, Matteo Turchetta, and Andreas Krause · 2018
Cited alongside, same era.
David Mguni · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Reward constrained policy optimization
Chen Tessler, Daniel J Mankowitz, and Shie Mannor · 2018
Cited alongside, same era.
Lyapunov-based safe policy optimization for continuous control
Yinlam Chow, Ofir Nachum, Aleksandra Faust, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2019
Cited alongside, same era.
Lyapunov-based safe policy optimization for continuous control, 2019
Yinlam Chow, Ofir Nachum, Aleksandra Faust, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2019
Cited alongside, same era.
Reinforcement learning with almost sure constraints
Agustin Castellano, Hancheng Min, Juan Bazerque, and Enrique Mallada · 2021
Closest in time.
On the complexity of computing markov perfect equilibrium in general-sum stochastic games
Xiaotie Deng, Yuhao Li, David Henry Mguni, Jun Wang, and Yaodong Yang · 2021
Closest in time.
Multi-agent constrained policy optimisation
Shangding Gu, Jakub Grudzien Kuba, Munning Wen, Ruiqing Chen, Ziyan Wang, Zheng Tian, Jun Wang, Alois Knoll, and Yaodong Yang · 2021
Closest in time.
Trust region policy optimisation in multi-agent reinforcement learning
Jakub Grudzien Kuba, Ruiqing Chen, Munning Wen, Ying Wen, Fanglei Sun, Jun Wang, and Yaodong Yang · 2021
Closest in time.
Ligs: Learnable intrinsic-reward generation selection for multi-agent learning
David Henry Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez-Nieves, Oliver Slumbers, Feifei Tong, Yang Li, Jiangcheng Zhu, Yaodong Yang, and Jun Wang · 2021
Closest in time.
Constrained policy optimization via Bayesian world models
Yarden As, Ilnura Usmanova, Sebastian Curi, and Andreas Krause · 2022
Closest in time.
Constrained variational policy optimization for safe reinforcement learning
Zuxin Liu, Zhepeng Cen, Vladislav Isenbaev, Wei Liu, Zhiwei Steven Wu, Bo Li, and Ding Zhao · 2022
Closest in time.
Timing is everything: Learning to act selectively with costly actions and budgetary constraints
David Mguni, Aivar Sootla, Juliusz Ziomek, Oliver Slumbers, Zipeng Dai, Kun Shao, and Jun Wang · 2022
Closest in time.
SAUTE RL: Almost surely safe reinforcement learning using state augmentation, 2022
Aivar Sootla, Alexander I. Cowen-Rivers, Taher Jafferjee, Ziyan Wang, David Mguni, Jun Wang, and Haitham Bou-Ammar · 2022
Closest in time.
Seren: Knowing when to explore and when to exploit
Changmin Yu, David Mguni, Dong Li, Aivar Sootla, Jun Wang, and Neil Burgess · 2022
Closest in time.