Fetching the paper…
Reading the bibliography…
Policy gradient (PG) methods are popular reinforcement learning (RL) methods where a baseline is often applied to reduce the variance of gradient estimates.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Reinforcement learning in pomdps with function approximation
Hajime Kimura, Kazuteru Miyazaki, and Shigenobu Kobayashi · 1997
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. Mcallester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
The optimal reward baseline for gradient-based reinforcement learning
Lex Weaver and Nigel Tao · 2001
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
Policy gradient methods for robotics
Jan Peters and Stefan Schaal · 2006
Earlier work this paper cites.
Analysis and improvement of policy gradient estimation
Tingting Zhao, Hirotaka Hachiya, Gang Niu, and Masashi Sugiyama · 2011
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Q-prop: Sample-efficient policy gradient with an off-policy critic
S Gu, T Lillicrap, Z Ghahramani, RE Turner, and S Levine · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
Multiagent bidirectionally-coordinated nets for learning to play starcraft combat games. arxiv 2017
P Peng, Q Yuan, Y Wen, Y Yang, Z Tang, H Long, and J Wang · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, F. Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Spinning Up in Deep Reinforcement Learning
Joshua Achiam · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Factorized q-learning for large-scale multi-agent systems
Ming Zhou, Yong Chen, Ying Wen, Yaodong Yang, Yufeng Su, Weinan Zhang, Dell Zhang, and Jun Wang · 2019
Later among the works it cites.
Is independent learning all you need in the starcraft multi-agent challenge?
Christian Schroeder de Witt, Tarun Gupta, Denys Makoviichuk, Viktor Makoviychuk, Philip HS Torr, Mingfei Sun, and Shimon Whiteson · 2020
Later among the works it cites.
Deep multi-agent reinforcement learning for decentralized continuous cooperative control
Christian Schröder de Witt, Bei Peng, Pierre-Alexandre Kamienny, Philip H. S. Torr, Wendelin Böhmer, and Shimon Whiteson · 2020
Later among the works it cites.
Expected policy gradients for reinforcement learning
Shimon Whiteson Kamil Ciosek · 2020
Later among the works it cites.
Comparative evaluation of cooperative multi-agent deep reinforcement learning algorithms
Georgios Papoudakis, Filippos Christianos, Lukas Schäfer, and Stefano V Albrecht · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
The mirage of action-dependent baselines in reinforcement learning
George Tucker, Surya Bhupatiraju, Shixiang Gu, Richard Turner, Zoubin Ghahramani, and Sergey Levine · 2018
Cited alongside, same era.
Probabilistic recursive reasoning for multi-agent reinforcement learning
Ying Wen, Yaodong Yang, Rui Luo, Jun Wang, and Wei Pan · 2018
Cited alongside, same era.
Variance reduction for policy gradient with action-dependent factorized baselines
Cathy Wu, Aravind Rajeswaran, Yan Duan, Vikash Kumar, Alexandre M Bayen, Sham Kakade, Igor Mordatch, and Pieter Abbeel · 2018
Cited alongside, same era.
Mean field multi-agent reinforcement learning
Yaodong Yang, Rui Luo, Minne Li, Ming Zhou, Weinan Zhang, and Jun Wang · 2018
Cited alongside, same era.
Later among the works it cites.
Facmac: Factored multi-agent centralised policy gradients
Bei Peng, Tabish Rashid, Christian A Schroeder de Witt, Pierre-Alexandre Kamienny, Philip HS Torr, Wendelin Böhmer, and Shimon Whiteson · 2020
Later among the works it cites.
Is independent learning all you need in the starcraft multi-agent challenge?
Christian Schroeder de Witt, Tarun Gupta, Denys Makoviichuk, Viktor Makoviychuk, Philip HS Torr, Mingfei Sun, and Shimon Whiteson · 2020
Later among the works it cites.
Off-policy multi-agent decomposed policy gradients
Yihan Wang, Beining Han, Tonghan Wang, Heng Dong, and Chongjie Zhang · 2020
Later among the works it cites.
Modelling bounded rationality in multi-agent interactions by generalized recursive reasoning
Ying Wen, Yaodong Yang, and Jun Wang · 2020
Later among the works it cites.
An overview of multi-agent reinforcement learning from game theoretical perspective
Yaodong Yang and Jun Wang · 2020
Later among the works it cites.
Multi-agent determinantal q-learning
Yaodong Yang, Ying Wen, Jun Wang, Liheng Chen, Kun Shao, David Mguni, and Weinan Zhang · 2020
Later among the works it cites.
Bi-level actor-critic for multi-agent coordination
Haifeng Zhang, Weizhe Chen, Zeren Huang, Minne Li, Yaodong Yang, Weinan Zhang, and Jun Wang · 2020
Later among the works it cites.
Smarts: Scalable multi-agent reinforcement learning training school for autonomous driving
Ming Zhou, Jun Luo, Julian Villella, Yaodong Yang, David Rusu, Jiayu Miao, Weinan Zhang, Montgomery Alban, Iman Fadakar, Zheng Chen, et al · 2020
Later among the works it cites.
Contrasting centralized and decentralized critics in multi-agent reinforcement learning
Xueguang Lyu, Yuchen Xiao, Brett Daley, and Christopher Amato · 2021
Closest in time.
Learning in nonzero-sum stochastic games with potentials
David H Mguni, Yutong Wu, Yali Du, Yaodong Yang, Ziyi Wang, Minne Li, Ying Wen, Joel Jennings, and Jun Wang · 2021
Closest in time.
The surprising effectiveness of mappo in cooperative, multi-agent games
Chao Yu, A. Velu, Eugene Vinitsky, Yu Wang, A. Bayen, and Yi Wu · 2021
Closest in time.