Fetching the paper…
Reading the bibliography…
Safe reinforcement learning (RL) is still very challenging since it requires the agent to consider both return maximization and safe exploration.
A markovian decision process
Bellman, R. (1957) · 1957
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H. (1989) · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Constrained Markov decision processes
Altman, E. (1999) · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Earlier work this paper cites.
An introduction to numerical analysis
Süli, E. and Mayers, D. F. (2003) · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Greensmith, E., Bartlett, P. L., and Baxter, J. (2004) · 2004
Earlier work this paper cites.
Accelerating safe reinforcement learning with constraint-mismatched policies
Yang, T.-Y., Rosca, J., Narasimhan, K., and Ramadge, P. J. (2020a) · 2006
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Peters, J. and Schaal, S. (2008) · 2008
Earlier work this paper cites.
Robust constrained-mdps: Soft-constrained robust policy optimization under model uncertainty
Russel, R. H., Benosman, M., and Van Baar, J. (2020) · 2010
Earlier work this paper cites.
Information theory: coding theorems for discrete memoryless systems
Csiszár, I. and Körner, J. (2011) · 2011
Earlier work this paper cites.
Han, M., Tian, Yuanand Zhang, L., Wang, J., and Pan, W. (2020) · 2011
Earlier work this paper cites.
A primal approach to constrained policy optimization: Global optimality and finite-time analysis
Xu, T., Liang, Y., and Lan, G. (2020) · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Earlier work this paper cites.
A survey on policy search for robotics
Deisenroth, M. P., Neumann, G., and Peters, J. (2013) · 2013
Earlier work this paper cites.
Safe policy iteration
Pirotta, M., Restelli, M., Pecorino, A., and Calandriello, D. (2013) · 2013
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
A unified approach for multi-step temporal-difference learning with eligibility traces in reinforcement learning
Yang, L., Shi, M., Zheng, Q., Meng, W., and Pan, G. (2018) · 2018
Later among the works it cites.
Batch policy learning under constraints
Le, H., Voloshin, C., and Yue, Y. (2019) · 2019
Later among the works it cites.
Openai five defeats dota 2 world champions
OpenAI (2019) · 2019
Later among the works it cites.
Constrained reinforcement learning has zero duality gap
Paternain, S., Chamon, L. F., Calvo-Fullana, M., and Ribeiro, A. (2019) · 2019
Later among the works it cites.
Benchmarking Safe Exploration in Deep Reinforcement Learning
Ray, A., Achiam, J., and Amodei, D. (2019) · 2019
Later among the works it cites.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Chow, Y., Ghavamzadeh, M., Janson, L., and Pavone, M. (2017) · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. (2017) · 2017
Cited alongside, same era.
Alphastar: Mastering the real-time strategy game starcraft ii
Vinyals, O., Babuschkin, I., Chung, J., Mathieu, M., Jaderberg, M., Czarnecki, W. M., Dudzik, A., Huang, A., Georgiev, P., Powell, R., et al. (2019) · 2019
Later among the works it cites.
Supervised policy update for deep reinforcement learning
Vuong, Q., Zhang, Y., and Ross, K. W. (2019) · 2019
Later among the works it cites.
Risk-averse trust region optimization for reward-volatility reduction
Bisi, L., Sabbioni, L., Vittori, E., Papini, M., and Restelli, M. (2020) · 2020
Later among the works it cites.
Constrained markov decision processes via backward value functions
Satija, H., Amortila, P., and Pineau, J. (2020) · 2020
Later among the works it cites.
First order constrained optimization in policy space
Zhang, Y., Vuong, Q., and Ross, K. (2020) · 2020
Later among the works it cites.
Reinforcement learning based recommender systems: A survey
Afsar, M. M., Crump, T., and Far, B. (2021) · 2021
Later among the works it cites.
Conservative safety critics for exploration
Bharadhwaj, H., Kumar, A., Rhinehart, N., Levine, S., Shkurti, F., and Garg, A. (2021) · 2021
Later among the works it cites.
A primal-dual approach to constrained markov decision processes
Chen, Y., Dong, J., and Wang, Z. (2021) · 2021
Later among the works it cites.
Learning safe policies with cost-sensitive advantage estimation
Kang, B., Mannor, S., and Feng, J. (2021) · 2021
Later among the works it cites.
Safe continuous control with constrained model-based policy optimization
Zanger, M. A., Daaboul, K., and Zöllner, J. M. (2021) · 2021
Later among the works it cites.