Fetching the paper…
Reading the bibliography…
Developing reinforcement learning algorithms that satisfy safety constraints is becoming increasingly important in real-world applications.
Constrained Markov decision processes , volume 7
Altman, E · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., Mansour, Y., et al · 1999
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Asymptopia: an exposition of statistical asymptotic theory. 2000
Pollard, D · 2000
Earlier work this paper cites.
Safe exploration in markov decision processes
Moldovan, T. M. and Abbeel, P · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Control barrier certificates for safe swarm behavior
Borrmann, U., Wang, L., Ames, A. D., and Egerstedt, M · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Control barrier function based quadratic programs for safety critical systems
Ames, A. D., Xu, X., Grizzle, J. W., and Tabuada, P · 2016
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S., Shammah, S., and Shashua, A · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Chow, Y., Ghavamzadeh, M., Janson, L., and Pavone, M · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
A study of ai population dynamics with million-agent reinforcement learning
Yang, Y., Yu, L., Bai, Y., Wang, J., Zhang, W., Wen, Y., and Yu, Y · 2017
Cited alongside, same era.
Is independent learning all you need in the starcraft multi-agent challenge?
de Witt, C. S., Gupta, T., Makoviichuk, D., Makoviychuk, V., Torr, P. H., Sun, M., and Whiteson, S · 2020
Later among the works it cites.
Facmac: Factored multi-agent centralised policy gradients
Peng, B., Rashid, T., de Witt, C. A. S., Kamienny, P.-A., Torr, P. H., Böhmer, W., and Whiteson, S · 2020
Later among the works it cites.
Learning safe multi-agent control with decentralized neural barrier certificates
Qin, Z., Zhang, K., Chen, Y., Chen, J., and Fan, C · 2020
Later among the works it cites.
An overview of multi-agent reinforcement learning from game theoretical perspective
Yang, Y. and Wang, J · 2020
Later among the works it cites.
Multi-agent determinantal q-learning
Yang, Y., Wen, Y., Wang, J., Chen, L., Shao, K., Mguni, D., and Zhang, W · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chow, Y., Nachum, O., Duenez-Guzman, E., and Ghavamzadeh, M · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., Schroeder, C., Farquhar, G., Foerster, J., and Whiteson, S · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., et al · 2018
Cited alongside, same era.
Effortless creation of safe robots from modules through self-programming and self-verification
Althoff, M., Giusti, A., Liu, S. B., and Pereira, A · 2019
Cited alongside, same era.
Lyapunov-based safe policy optimization for continuous control
Chow, Y., Nachum, O., Faust, A., Duenez-Guzman, E., and Ghavamzadeh, M · 2019
Cited alongside, same era.
Benchmarking safe exploration in deep reinforcement learning
Ray, A., Achiam, J., and Amodei, D · 2019
Cited alongside, same era.
Later among the works it cites.
Smarts: Scalable multi-agent reinforcement learning training school for autonomous driving
Zhou, M., Luo, J., Villela, J., Yang, Y., Rusu, D., Miao, J., Zhang, W., Alban, M., Fadakar, I., Chen, Z., et al · 2020
Later among the works it cites.
robosuite: A modular simulation framework and benchmark for robot learning
Zhu, Y., Wong, J., Mandlekar, A., and Martín-Martín, R · 2020
Later among the works it cites.
On the complexity of computing markov perfect equilibrium in general-sum stochastic games
Deng, X., Li, Y., Mguni, D. H., Wang, J., and Yang, Y · 2021
Closest in time.
Cmix: Deep multi-agent reinforcement learning with peak and average constraints
Liu, C., Geng, N., Aggarwal, V., Lan, T., Yang, Y., and Xu, M · 2021
Closest in time.
Decentralized policy gradient descent ascent for safe multi-agent reinforcement learning
Lu, S., Zhang, K., Chen, T., Basar, T., and Horesh, L · 2021
Closest in time.
A provably-efficient model-free algorithm for constrained markov decision processes
Wei, H., Liu, X., and Ying, L · 2021
Closest in time.
Crpo: A new approach for safe reinforcement learning with convergence guarantee
Xu, T., Liang, Y., and Lan, G · 2021
Closest in time.
The surprising effectiveness of mappo in cooperative, multi-agent games
Yu, C., Velu, A., Vinitsky, E., Wang, Y., Bayen, A., and Wu, Y · 2021
Closest in time.
Safe continuous control with constrained model-based policy optimization, 2021
Zanger, M. A., Daaboul, K., and Zöllner, J. M · 2021
Closest in time.