Fetching the paper…
Reading the bibliography…
For many applications of reinforcement learning it can be more convenient to specify both a reward function and constraints, rather than trying to design behavior through the reward function.
Information Theory: Coding Theorems for Discrete Memoryless Systems
Csiszar, I and Körner, J · 1981
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
Constrained Markov Decision Processes
Altman, Eitan · 1999
Earlier work this paper cites.
Policy invariance under reward transformations : Theory and application to reward shaping
Ng, Andrew Y., Harada, Daishi, and Russell, Stuart · 1999
Earlier work this paper cites.
Approximately Optimal Approximate Reinforcement Learning
Kakade, Sham and Langford, John · 2002
Earlier work this paper cites.
Subgradient methods
Boyd, Stephen, Xiao, Lin, and Mutapcic, Almir · 2003
Earlier work this paper cites.
Constrained reinforcement learning from intrinsic and extrinsic rewards
Uchibe, Eiji and Doya, Kenji · 2007
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Peters, Jan and Schaal, Stefan · 2008
Earlier work this paper cites.
Safe Exploration in Markov Decision Processes
Moldovan, Teodor Mihai and Abbeel, Pieter · 2012
Earlier work this paper cites.
Safe Policy Iteration
Pirotta, Matteo, Restelli, Marcello, and Calandriello, Daniele · 2013
Cited alongside, same era.
Safe Policy Search for Lifelong Reinforcement Learning with Sublinear Regret
Bou Ammar, Haitham, Tutunov, Rasul, and Eaton, Eric · 2015
Cited alongside, same era.
Risk-Constrained Reinforcement Learning with Percentile Risk Criteria
Chow, Yinlam, Ghavamzadeh, Mohammad, Janson, Lucas, and Pavone, Marco · 2015
Cited alongside, same era.
A Comprehensive Survey on Safe Reinforcement Learning
García, Javier and Fernández, Fernando · 2015
Cited alongside, same era.
Doubly Robust Off-policy Value Evaluation for Reinforcement Learning
Jiang, Nan and Li, Lihong · 2015
Cited alongside, same era.
End-to-End Training of Deep Visuomotor Policies
Levine, Sergey, Finn, Chelsea, Darrell, Trevor, and Abbeel, Pieter · 2016
Later among the works it cites.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P., Hunt, Jonathan J., Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2016
Later among the works it cites.
Asynchronous Methods for Deep Reinforcement Learning
Mnih, Volodymyr, Badia, Adrià Puigdomènech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy P., Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
High-Dimensional Continuous Control Using Generalized Advantage Estimation
Schulman, John, Moritz, Philipp, Levine, Sergey, Jordan, Michael, and Abbeel, Pieter · 2016
Later among the works it cites.
Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving
Shalev-Shwartz, Shai, Shammah, Shaked, and Shashua, Amnon · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei a, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Cited alongside, same era.
Trust Region Policy Optimization
Schulman, John, Moritz, Philipp, Jordan, Michael, and Abbeel, Pieter · 2015
Cited alongside, same era.
Concrete Problems in AI Safety
Amodei, Dario, Olah, Chris, Steinhardt, Jacob, Christiano, Paul, Schulman, John, and Mané, Dan · 2016
Cited alongside, same era.
Benchmarking Deep Reinforcement Learning for Continuous Control
Duan, Yan, Chen, Xi, Schulman, John, and Abbeel, Pieter · 2016
Cited alongside, same era.
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
Silver, David, Huang, Aja, Maddison, Chris J., Guez, Arthur, Sifre, Laurent, van den Driessche, George, Schrittwieser, Julian, Antonoglou, Ioannis, Panneershelvam, Veda, Lanctot, Marc, Dieleman, Sander, Grewe, Dominik, Nham, John, Kalchbrenner, Nal, Sutskever, Ilya, Lillicrap, Timothy, Leach, Madeleine, Kavukcuoglu, Koray, Graepel, Thore, and Hassabis, Demis · 2016
Later among the works it cites.
Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic
Gu, Shixiang, Lillicrap, Timothy, Ghahramani, Zoubin, Turner, Richard E., and Levine, Sergey · 2017
Closest in time.
Probabilistically Safe Policy Transfer
Held, David, Mccarthy, Zoe, Zhang, Michael, Shentu, Fred, and Abbeel, Pieter · 2017
Closest in time.
Combating Deep Reinforcement Learning’s Sisyphean Curse with Intrinsic Fear
Lipton, Zachary C., Gao, Jianfeng, Li, Lihong, Chen, Jianshu, and Deng, Li · 2017
Closest in time.