Fetching the paper…
Reading the bibliography…
Safe reinforcement learning (RL) studies problems where an intelligent agent has to not only maximize reward but also avoid exploring unsafe areas.
A markovian decision process
Richard Bellman · 1957
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Nonlinear programming
Dimitri P Bertsekas · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Constrained Markov decision processes
Eitan Altman · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
An introduction to numerical analysis
Endre Süli and David F Mayers · 2003
Earlier work this paper cites.
Convex optimization
Stephen P Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
A tutorial on mm algorithms
David R Hunter and Kenneth Lange · 2004
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Jan. Peters and Stefan. Schaal · 2008
Earlier work this paper cites.
Information theory: coding theorems for discrete memoryless systems
Imre Csiszár and János Körner · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Safe policy iteration
M. Pirotta, M. Restelli, A. Pecorino, and D. Calandriello · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Cited alongside, same era.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Benchmarking Safe Exploration in Deep Reinforcement Learning
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Later among the works it cites.
Reward constrained policy optimization
Chen Tessler, Daniel J Mankowitz, and Shie Mannor · 2019
Later among the works it cites.
Supervised policy update for deep reinforcement learning
Quan Vuong, Yiming Zhang, and Keith W Ross · 2019
Later among the works it cites.
Risk-averse trust region optimization for reward-volatility reduction
Lorenzo Bisi, Luca Sabbioni, Edoardo Vittori, Matteo Papini, and Marcello Restelli · 2020
Later among the works it cites.
Minghao Han, Lixian Tian, Yuanand Zhang, Jun Wang, and Wei Pan · 2020
Later among the works it cites.
Ipo: Interior-point policy optimization under constraints
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
A lyapunov-based approach to safe reinforcement learning
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2018
Cited alongside, same era.
Safe exploration in continuous action spaces
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa · 2018
Cited alongside, same era.
Accelerated primal-dual policy optimization for safe reinforcement learning
Qingkai Liang, Fanyu Que, and Eytan Modiano · 2018
Cited alongside, same era.
Variance reduction for policy gradient with action-dependent factorized baselines
Cathy Wu, Aravind Rajeswaran, Yan Duan, Vikash Kumar, Alexandre M Bayen, Sham Kakade, Igor Mordatch, and Pieter Abbeel · 2018
Cited alongside, same era.
A unified approach for multi-step temporal-difference learning with eligibility traces in reinforcement learning
Long Yang, Minhao Shi, Qian Zheng, Wenjia Meng, and Gang Pan · 2018
Cited alongside, same era.
Yongshuai Liu, Jiaxin Ding, and Xin Liu · 2020
Later among the works it cites.
Robust constrained-mdps: Soft-constrained robust policy optimization under model uncertainty
Reazul Hasan Russel, Mouhacine Benosman, and Jeroen Van Baar · 2020
Later among the works it cites.
Constrained markov decision processes via backward value functions
Harsh Satija, Philip Amortila, and Joelle Pineau · 2020
Later among the works it cites.
Projection-based constrained policy optimization
Tsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, and Peter J Ramadge · 2020
Later among the works it cites.
First order constrained optimization in policy space
Yiming Zhang, Quan Vuong, and Keith Ross · 2020
Later among the works it cites.
Reinforcement learning based recommender systems: A survey
M Mehdi Afsar, Trafford Crump, and Behrouz Far · 2021
Later among the works it cites.
Conservative safety critics for exploration
Homanga Bharadhwaj, Aviral Kumar, Nicholas Rhinehart, Sergey Levine, Florian Shkurti, and Animesh Garg · 2021
Later among the works it cites.
A primal-dual approach to constrained markov decision processes
Yi Chen, Jing Dong, and Zhaoran Wang · 2021
Later among the works it cites.
Learning safe policies with cost-sensitive advantage estimation, 2021
Bingyi Kang, Shie Mannor, and Jiashi Feng · 2021
Later among the works it cites.
On convergence of gradient expected sarsa ( λ \lambda )
Long Yang, Gang Zheng, Yu Zhang, Qian Zheng, Pengfei Li, and Gang Pan · 2021
Later among the works it cites.
Sample complexity of policy gradient finding second-order stationary points
Long Yang, Qian Zheng, and Gang Pan · 2021
Later among the works it cites.
Policy optimization with stochastic mirror descent
Long Yang, Gang Zheng, Haotian Zhang, Yu Zhang, Qian Zheng, and Gang Pan · 2022
Closest in time.