Fetching the paper…
Reading the bibliography…
We investigate safe multi-agent reinforcement learning, where agents seek to collectively maximize an aggregate sum of local objectives while satisfying their own safety constraints.
Some np-complete problems in quadratic and nonlinear programming
Katta G Murty and Santosh N Kabadi · 1987
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Nonlinear programming
Dimitri P Bertsekas · 1997
Earlier work this paper cites.
Constrained Markov decision processes
Eitan Altman · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
A survey of computational complexity results in systems and control
Vincent D Blondel and John N Tsitsiklis · 2000
Earlier work this paper cites.
Trust region methods
Andrew R Conn, Nicholas IM Gould, and Philippe L Toint · 2000
Earlier work this paper cites.
Constrained markov games: Nash equilibria
Eitan Altman and Adam Shwartz · 2000
Earlier work this paper cites.
Design challenges for energy-constrained ad hoc wireless networks
Andrea J Goldsmith and Stephen B Wicker · 2002
Earlier work this paper cites.
Saddle-point calculation for constrained finite markov chains
E Gómez-Ramırez, K Najim, and AS Poznyak · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe · 2004
Earlier work this paper cites.
A characterization of convex problems in decentralized control
Michael Rotkowitz and Sanjay Lall · 2005
Earlier work this paper cites.
Fundamentals of wireless communication
David Tse and Pramod Viswanath · 2005
Earlier work this paper cites.
Constrained cost-coupled stochastic games with independent state processes
Eitan Altman, Konstantin Avrachenkov, Nicolas Bonneau, Merouane Debbah, Rachid El-Azouzi, and Daniel Sadoc Menasche · 2008
Earlier work this paper cites.
Influence maximization in social networks when negative opinions may emerge and propagate
Wei Chen, Alex Collins, Rachel Cummings, Te Ke, Zhenming Liu, David Rincon, Xiaorui Sun, Yajun Wang, Wei Wei, and Yifei Yuan · 2011
Earlier work this paper cites.
Gibbs measures and phase transitions
Hans-Otto Georgii · 2011
Earlier work this paper cites.
The effect of the interbank network structure on contagion and common shocks
Co-Pierre Georg · 2013
Earlier work this paper cites.
Platoon-based multi-agent intersection management for connected vehicle
Qiu Jin, Guoyuan Wu, Kanok Boriboonsomsin, and Matthew Barth · 2013
Earlier work this paper cites.
Correlation decay method for decision, optimization, and inference in large-scale networks
David Gamarnik · 2013
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Earlier work this paper cites.
A characterization of stationary nash equilibria of constrained stochastic games with independent state processes
Vikas Vikram Singh and N Hemachandra · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Optimal resource allocation for control of networked epidemic models
Cameron Nowzari, Victor M Preciado, and George J Pappas · 2015
Earlier work this paper cites.
Necessary and sufficient conditions for optimality in constrained general sum stochastic games
Vinayaka G Yaji and Shalabh Bhatnagar · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Earlier work this paper cites.
Sub-sampled cubic regularization for non-convex optimization
Jonas Moritz Kohler and Aurelien Lucchi · 2017
Earlier work this paper cites.
Wasserstein generative adversarial networks
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2018
Cited alongside, same era.
Networking and communications in autonomous driving: A survey
Jiadai Wang, Jiajia Liu, and Nei Kato · 2018
Cited alongside, same era.
Safety-aware apprenticeship learning
Weichao Zhou and Wenchao Li · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Concentration bounds for two time scale stochastic approximation
Vivek S Borkar and Sarath Pattathil · 2018
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Decentralized policy gradient descent ascent for safe multi-agent reinforcement learning
Songtao Lu, Kaiqing Zhang, Tianyi Chen, Tamer Başar, and Lior Horesh · 2021
Later among the works it cites.
Multi-agent constrained policy optimisation
Shangding Gu, Jakub Grudzien Kuba, Munning Wen, Ruiqing Chen, Ziyan Wang, Zheng Tian, Jun Wang, Alois Knoll, and Yaodong Yang · 2021
Later among the works it cites.
Dimension-free rates for natural policy gradient in multi-agent reinforcement learning
Carlo Alfano and Patrick Rebeschini · 2021
Later among the works it cites.
Multi-agent reinforcement learning in stochastic networked systems
Yiheng Lin, Guannan Qu, Longbo Huang, and Adam Wierman · 2021
Later among the works it cites.
Pettingzoo: Gym for multi-agent reinforcement learning
J Terry, Benjamin Black, Nathaniel Grammel, Mario Jayakumar, Ananth Hari, Ryan Sullivan, Luis S Santos, Clemens Dieffendahl, Caroline Horsch, Rodrigo Perez-Vicente, et al · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Cited alongside, same era.
Efficient exploration via state marginal matching
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Cited alongside, same era.
Solving a class of non-convex min-max games using iterative first order methods
Maher Nouiehed, Maziar Sanjabi, Tianjian Huang, Jason D Lee, and Meisam Razaviyayn · 2019
Cited alongside, same era.
Safe policies for reinforcement learning via primal-dual methods
Santiago Paternain, Miguel Calvo-Fullana, Luiz FO Chamon, and Alejandro Ribeiro · 2019
Cited alongside, same era.
Convergent policy optimization for safe reinforcement learning
Ming Yu, Zhuoran Yang, Mladen Kolar, and Zhaoran Wang · 2019
Cited alongside, same era.
Later among the works it cites.
Constrained multiagent markov decision processes: A taxonomy of problems and algorithms
Frits De Nijs, Erwin Walraven, Mathijs De Weerdt, and Matthijs Spaan · 2021
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo Jovanovic · 2021
Later among the works it cites.
Learning policies with zero or bounded constraint violation for constrained MDPs
Tao Liu, Ruida Zhou, Dileep Kalathil, PR Kumar, and Chao Tian · 2021
Later among the works it cites.
Constrained expected average stochastic games for continuous-time jump processes
Qingda Wei · 2021
Later among the works it cites.
On the convergence and sample efficiency of variance-reduced policy gradient method
Junyu Zhang, Chengzhuo Ni, Csaba Szepesvari, Mengdi Wang, et al · 2021
Later among the works it cites.
Concave utility reinforcement learning: the mean-field game viewpoint
Matthieu Geist, Julien Pérolat, Mathieu Laurière, Romuald Elie, Sarah Perrin, Olivier Bachem, Rémi Munos, and Olivier Pietquin · 2021
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2021
Later among the works it cites.
Sample complexity bounds for two timescale value-based reinforcement learning algorithms
Tengyu Xu and Yingbin Liang · 2021
Later among the works it cites.
Marl with general utilities via decentralized shadow reward actor-critic
Junyu Zhang, Amrit Singh Bedi, Mengdi Wang, and Alec Koppel · 2022
Later among the works it cites.
Scalable reinforcement learning for multiagent networked systems
Guannan Qu, Adam Wierman, and Na Li · 2022
Later among the works it cites.
Mean-field approximation of cooperative constrained multi-agent reinforcement learning (cmarl)
Washim Uddin Mondal, Vaneet Aggarwal, and Satish V Ukkusuri · 2022
Later among the works it cites.
Global convergence of localized policy iteration in networked multi-agent reinforcement learning
Yizhou Zhang, Guannan Qu, Pan Xu, Yiheng Lin, Zaiwei Chen, and Adam Wierman · 2022
Later among the works it cites.
Near-optimal distributed linear-quadratic regulator for networked systems
Sungho Shin, Yiheng Lin, Guannan Qu, Adam Wierman, and Mihai Anitescu · 2022
Later among the works it cites.
Xin Liu, Honghao Wei, and Lei Ying · 2022
Later among the works it cites.
A review of safe reinforcement learning: Methods, theory and applications
Shangding Gu, Long Yang, Yali Du, Guang Chen, Florian Walter, Jun Wang, Yaodong Yang, and Alois Knoll · 2022
Later among the works it cites.
A dual approach to constrained markov decision processes with entropy regularization
Donghao Ying, Yuhao Ding, and Javad Lavaei · 2022
Later among the works it cites.
Policy-based primal-dual methods for convex constrained markov decision processes
Donghao Ying, Mengzi Guo, Yuhao Ding, Javad Lavaei, et al · 2022
Later among the works it cites.
Yuhao Ding and Javad Lavaei · 2022
Later among the works it cites.
Achieving zero constraint violation for constrained reinforcement learning via primal-dual approach
Qinbo Bai, Amrit Singh Bedi, Mridul Agarwal, Alec Koppel, and Vaneet Aggarwal · 2022
Later among the works it cites.
Constrained average stochastic games with continuous-time independent state processes
Wenzhao Zhang and Xiaolong Zou · 2022
Later among the works it cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2022
Later among the works it cites.
Provably efficient generalized lagrangian policy optimization for safe multi-agent reinforcement learning
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo R Jovanovic · 2023
Closest in time.
Scalable multi-agent reinforcement learning with general utilities
Donghao Ying, Yuhao Ding, Alec Koppel, and Javad Lavaei · 2023
Closest in time.
Cem: Constrained entropy maximization for task-agnostic safe exploration
Qisong Yang and Matthijs TJ Spaan · 2023
Closest in time.