Fetching the paper…
Reading the bibliography…
We present a multi-agent actor-critic method that aims to implicitly address the credit assignment problem under fully cooperative settings.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
Advances in genetic programming
Kenneth E Kinnear, William B Langdon, Lee Spector, Peter J Angeline, and Una-May O’Reilly · 1994
Earlier work this paper cites.
Learning without state-estimation in partially observable markovian decision processes
Satinder P Singh, Tommi Jaakkola, and Michael I Jordan · 1994
Earlier work this paper cites.
Reinforcement learning algorithm for partially observable markov decision problems
Tommi Jaakkola, Satinder P Singh, and Michael I Jordan · 1995
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
Distributed agent-based air traffic flow management
Kagan Tumer and Adrian Agogino · 2007
Earlier work this paper cites.
Policy gradient critics
Daan Wierstra and Jürgen Schmidhuber · 2007
Earlier work this paper cites.
Reinforcement learning by value gradients
Michael Fairbank · 2008
Earlier work this paper cites.
Optimal and approximate q-value functions for decentralized pomdps
Frans A Oliehoek, Matthijs TJ Spaan, and Nikos Vlassis · 2008
Earlier work this paper cites.
An overview of recent progress in the study of distributed multi-agent coordination
Yongcan Cao, Wenwu Yu, Wei Ren, and Guanrong Chen · 2012
Earlier work this paper cites.
Value-gradient learning
Michael Fairbank and Eduardo Alonso · 2012
Earlier work this paper cites.
Modeling difference rewards for multiagent learning
Scott Proper and Kagan Tumer · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Weighted importance sampling for off-policy learning with linear function approximation
A Rupam Mahmood, Hado P van Hasselt, and Richard S Sutton · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Compatible value gradients for reinforcement learning of continuous deep policies
David Balduzzi and Muhammad Ghifary · 2015
Earlier work this paper cites.
Memory-based control with recurrent neural networks
Nicolas Heess, Jonathan J Hunt, Timothy P Lillicrap, and David Silver · 2015
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Bridging the gap between value and policy based reinforcement learning
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans · 2017
Later among the works it cites.
Equivalence between policy gradients and soft q-learning
John Schulman, Xi Chen, and Pieter Abbeel · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Later among the works it cites.
Counterfactual multi-agent policy gradients
Jakob N Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
A multi-agent framework for packet routing in wireless sensor networks
Dayong Ye, Minjie Zhang, and Yun Yang · 2015
Cited alongside, same era.
David Ha, Andrew Dai, and Quoc V Le · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Cited alongside, same era.
Multi-agent reinforcement learning as a rehearsal for decentralized planning
Landon Kraemer and Bikramjit Banerjee · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
A concise introduction to decentralized POMDPs
Frans A Oliehoek, Christopher Amato, et al · 2016
Cited alongside, same era.
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Qmix: monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder De Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Later among the works it cites.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2018
Later among the works it cites.
Understanding the impact of entropy on policy optimization
Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, and Dale Schuurmans · 2019
Later among the works it cites.
Liir: Learning individual intrinsic reward in multi-agent reinforcement learning
Yali Du, Lei Han, Meng Fang, Ji Liu, Tianhong Dai, and Dacheng Tao · 2019
Later among the works it cites.
If maxent rl is the answer, what is the question?
Benjamin Eysenbach and Sergey Levine · 2019
Later among the works it cites.
Actor-attention-critic for multi-agent reinforcement learning
Shariq Iqbal and Fei Sha · 2019
Later among the works it cites.
Maven: Multi-agent variational exploration
Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, and Shimon Whiteson · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
The StarCraft Multi-Agent Challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philiph H. S. Torr, Jakob Foerster, and Shimon Whiteson · 2019
Later among the works it cites.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi · 2019
Later among the works it cites.
Credit assignment techniques in stochastic computation graphs
Théophane Weber, Nicolas Heess, Lars Buesing, and David Silver · 2019
Later among the works it cites.
Shapley Q-value: A Local Reward Approach to Solve Global Reward Games
Jianhong Wang, Yuan Zhang, Tae-Kyun Kim, and Yunjie Gu · 2020
Closest in time.