Fetching the paper…
Reading the bibliography…
In this work, we consider the problem of computing optimal actions for Reinforcement Learning (RL) agents in a co-operative setting, where the objective is to optimize a common goal.
Stochastic approximation with two time scales
Vivek S Borkar · 1997
Earlier work this paper cites.
Inventory management in supply chains: a reinforcement learning approach
Ilaria Giannoccaro and Pierpaolo Pontrandolfo · 2002
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Zero-sum constrained stochastic games with independent state processes
Eitan Altman, Konstantin Avrachenkov, Richard Marquez, and Gregory Miller · 2005
Earlier work this paper cites.
An actor-critic algorithm for constrained markov decision processes
Vivek S Borkar · 2005
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter · 2008
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint
Vivek S Borkar · 2009
Earlier work this paper cites.
An actor–critic algorithm with function approximation for discounted cost constrained markov decision processes
Shalabh Bhatnagar · 2010
Earlier work this paper cites.
Resource allocation among agents with mdp-induced preferences
Dmitri A Dolgov and Edmund H Durfee · 2011
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained markov decision processes
Shalabh Bhatnagar and K Lakshmanan · 2012
Earlier work this paper cites.
Game-theoretic methods for the smart grid: An overview of microgrid systems, demand-side management, and smart grid communications
Walid Saad, Zhu Han, H Vincent Poor, and Tamer Basar · 2012
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Scalable greedy algorithms for task/resource constrained multi-agent stochastic planning
Pritee Agrawal, Pradeep Varakantham, and William Yeoh · 2016
Cited alongside, same era.
Budget allocation using weakly coupled, constrained markov decision processes
Craig Boutilier and Tyler Lu · 2016
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Cited alongside, same era.
Decision-making policies for heterogeneous autonomous multi-agent systems with safety constraints
Ruohan Zhang, Yue Yu, Mahmoud El Chamie, Behçet Açikmese, and Dana H Ballard · 2016
Cited alongside, same era.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Accelerated primal-dual policy optimization for safe reinforcement learning
Qingkai Liang, Fanyu Que, and Eytan Modiano · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Reward constrained policy optimization
Chen Tessler, Daniel J Mankowitz, and Shie Mannor · 2018
Later among the works it cites.
Gang Chen · 2019
Later among the works it cites.
Actor-critic algorithms for constrained multi-agent reinforcement learning
Raghuram Bharadwaj Diddigi, Sai Koti Reddy Danda, Prabuchandran K.J., and Shalabh Bhatnagar · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Multiagent cooperation and competition with deep reinforcement learning
Ardi Tampuu, Tambet Matiisen, Dorian Kodelja, Ilya Kuzovkin, Kristjan Korjus, Juhan Aru, Jaan Aru, and Raul Vicente · 2017
Cited alongside, same era.
Constrained-action pomdps for multi-agent intelligent knowledge distribution
Michael Fowler, Pratap Tokekar, T Charles Clancy, and Ryan K Williams · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Learning attentional communication for multi-agent cooperation
Jiechuan Jiang and Zongqing Lu · 2018
Cited alongside, same era.
Later among the works it cites.
Actor-critic algorithms for constrained multi-agent reinforcement learning
Raghuram Bharadwaj Diddigi, D Reddy, Prabuchandran KJ, and Shalabh Bhatnagar · 2019
Later among the works it cites.
Actor-attention-critic for multi-agent reinforcement learning
Shariq Iqbal and Fei Sha · 2019
Later among the works it cites.
Modelling the dynamic joint policy of teammates with attention multi-agent ddpg
Hangyu Mao, Zhengchao Zhang, Zhen Xiao, and Zhibo Gong · 2019
Later among the works it cites.
A review of cooperative multi-agent deep reinforcement learning
Afshin OroojlooyJadid and Davood Hajinezhad · 2019
Later among the works it cites.
Risk averse reinforcement learning for mixed multi-agent environments
D Sai Koti Reddy, Amrita Saha, Srikanth G Tamilselvam, Priyanka Agrawal, and Pankaj Dayama · 2019
Later among the works it cites.
Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications
Thanh Thi Nguyen, Ngoc Duy Nguyen, and Saeid Nahavandi · 2020
Later among the works it cites.