Fetching the paper…
Reading the bibliography…
Value factorization is a popular and promising approach to scaling up multi-agent reinforcement learning in cooperative settings, which balances the learning scalability and the representational capacity of value functions.
Über abbildung von mannigfaltigkeiten
Luitzen Egbertus Jan Brouwer · 1911
Earlier work this paper cites.
On the reciprocal of the general algebraic matrix
Eliakim H Moore · 1920
Earlier work this paper cites.
Memoryless policies: theoretical limitations and practical results
Michael L Littman · 1994
Earlier work this paper cites.
Stable function approximation in dynamic programming
Geoffrey J Gordon · 1995
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes
Craig Boutilier · 1996
Earlier work this paper cites.
Feature-based methods for large scale dynamic programming
John N Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Richard S Sutton · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Learning finite-state controllers for partially observable environments
Nicolas Meuleau, Leonid Peshkin, Kee-Eung Kim, and Leslie Pack Kaelbling · 1999
Earlier work this paper cites.
On the undecidability of probabilistic planning and infinite-horizon partially observable markov decision problems
Omid Madani, Steve Hanks, and Anne Condon · 1999
Earlier work this paper cites.
Optimal payoff functions for members of collectives
David H Wolpert and Kagan Tumer · 2002
Earlier work this paper cites.
Cooperative multi-agent learning: The state of the art
Liviu Panait and Sean Luke · 2005
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Optimal and approximate q-value functions for decentralized pomdps
Frans A Oliehoek, Matthijs TJ Spaan, and Nikos Vlassis · 2008
Earlier work this paper cites.
Analyzing and visualizing multiagent rewards in dynamic and stochastic domains
Adrian K Agogino and Kagan Tumer · 2008
Earlier work this paper cites.
Multi-agent learning with policy prediction
Chongjie Zhang and Victor Lesser · 2010
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Amir-massoud Farahmand, Csaba Szepesvári, and Rémi Munos · 2010
Earlier work this paper cites.
Coordinated multi-agent reinforcement learning in networked distributed pomdps
Chongjie Zhang and Victor Lesser · 2011
Earlier work this paper cites.
An overview of recent progress in the study of distributed multi-agent coordination
Yongcan Cao, Wenwu Yu, Wei Ren, and Guanrong Chen · 2012
Earlier work this paper cites.
Near-optimal reinforcement learning in factored mdps
Ian Osband and Benjamin Van Roy · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando De Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
Multi-agent reinforcement learning as a rehearsal for decentralized planning
Landon Kraemer and Bikramjit Banerjee · 2016
Cited alongside, same era.
Pac reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Cited alongside, same era.
Reinforcement learning in rich-observation mdps using spectral methods
Kamyar Azizzadenesheli, Alessandro Lazaric, and Animashree Anandkumar · 2016
Cited alongside, same era.
A concise introduction to decentralized POMDPs
Frans A Oliehoek, Christopher Amato, et al · 2016
Cited alongside, same era.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Later among the works it cites.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Max Simchowitz and Kevin G Jamieson · 2019
Later among the works it cites.
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, and Sham M Kakade · 2019
Later among the works it cites.
Influence-based multi-agent exploration
Tonghan Wang, Jianhao Wang, Yi Wu, and Chongjie Zhang · 2020
Closest in time.
Learning nearly decomposable value functions via communication minimization
Tonghan Wang, Jianhao Wang, Chongyi Zheng, and Chongjie Zhang · 2020
Closest in time.
Emergent tool use from multi-agent autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maximilian Hüttenrauch, Adrian Šošić, and Gerhard Neumann · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
Credit assignment for collective multiagent rl with global rewards
Duc Thien Nguyen, Akshat Kumar, and Hoong Chuin Lau · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Jakob N Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Cited alongside, same era.
Closest in time.
Multi-agent reinforcement learning with emergent roles
Tonghan Wang, Heng Dong, Victor Lesser, and Chongjie Zhang · 2020
Closest in time.
Qatten: A general framework for cooperative multiagent reinforcement learning
Yaodong Yang, Jianye Hao, Ben Liao, Kun Shao, Guangyong Chen, Wulong Liu, and Hongyao Tang · 2020
Closest in time.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Closest in time.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2020
Closest in time.
Weighted qmix: Expanding monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Gregory Farquhar, Bei Peng, and Shimon Whiteson · 2020
Closest in time.
Ai-qmix: attention and imagination for dynamic multi-agent reinforcement learning
Shariq Iqbal, Christian A Schroeder de Witt, Bei Peng, Wendelin Böhmer, Shimon Whiteson, and Fei Sha · 2020
Closest in time.
Deep coordination graphs
Wendelin Böhmer, Vitaly Kurin, and Shimon Whiteson · 2020
Closest in time.
Facmac: Factored multi-agent centralised policy gradients
Bei Peng, Tabish Rashid, Christian A Schroeder de Witt, Pierre-Alexandre Kamienny, Philip HS Torr, Wendelin Böhmer, and Shimon Whiteson · 2020
Closest in time.
Reinforcement learning in factored mdps: Oracle-efficient algorithms and tighter regret bounds for the non-episodic setting
Ziping Xu and Ambuj Tewari · 2020
Closest in time.
Qplex: Duplex dueling multi-agent q-learning
Jianhao Wang, Zhizhou Ren, Terry Liu, Yu Yang, and Chongjie Zhang · 2021
Closest in time.
Dop: Off-policy multi-agent decomposed policy gradients
Yihan Wang, Beining Han, Tonghan Wang, Heng Dong, and Chongjie Zhang · 2021
Closest in time.
Tesseract: Tensorised actors for multi-agent reinforcement learning
Anuj Mahajan, Mikayel Samvelyan, Lei Mao, Viktor Makoviychuk, Animesh Garg, Jean Kossaifi, Shimon Whiteson, Yuke Zhu, and Animashree Anandkumar · 2021
Closest in time.
Wei-Fang Sun, Cheng-Kuang Lee, and Chun-Yi Lee · 2021
Closest in time.
Softmax with regularization: Better value estimation in multi-agent reinforcement learning
Ling Pan, Tabish Rashid, Bei Peng, Longbo Huang, and Shimon Whiteson · 2021
Closest in time.
Fop: Factorizing optimal joint policy of maximum-entropy multi-agent reinforcement learning
Tianhao Zhang, Yueheng Li, Chen Wang, Guangming Xie, and Zongqing Lu · 2021
Closest in time.
Efficient reinforcement learning in factored mdps with application to constrained rl
Xiaoyu Chen, Jiachen Hu, Lihong Li, and Liwei Wang · 2021
Closest in time.