Fetching the paper…
Reading the bibliography…
Centralized Training with Decentralized Execution (CTDE) has been proven to be an effective paradigm in cooperative multi-agent reinforcement learning (MARL).
Sequence Modeling of Temporal Credit Assignment for Episodic Reinforcement Learning
Liu, Y.; Luo, Y.; Zhong, Y.; Chen, X.; Liu, Q.; and Peng, J. 2019a · 1905
Earlier work this paper cites.
Sequence modeling of temporal credit assignment for episodic reinforcement learning
Liu, Y.; Luo, Y.; Zhong, Y.; Chen, X.; Liu, Q.; and Peng, J. 2019b · 1905
Earlier work this paper cites.
A value for n-person games
Shapley, L. S. 1997 · 1997
Earlier work this paper cites.
Theory and application to reward shaping
Ng, A. Y.; Harada, D.; and Russell, S. 1999 · 1999
Earlier work this paper cites.
The Shapley value on convex geometries
Bilbao, J. M.; and Edelman, P. H. 2000 · 2000
Earlier work this paper cites.
Qatten: A General Framework for Cooperative Multiagent Reinforcement Learning
Yang, Y.; Hao, J.; Liao, B.; Shao, K.; Chen, G.; Liu, W.; and Tang, H. 2020 · 2002
Earlier work this paper cites.
Cooperative multi-agent learning: The state of the art
Panait, L.; and Luke, S. 2005 · 2005
Earlier work this paper cites.
A Comprehensive Survey of Multiagent Reinforcement Learning
Busoniu, L.; Babuska, R.; and De Schutter, B. 2008 · 2008
Earlier work this paper cites.
Bounding the estimation error of sampling-based Shapley value approximation
Maleki, S.; Tran-Thanh, L.; Hines, G.; Rahwan, T.; and Rogers, A. 2013 · 2013
Earlier work this paper cites.
Long short-term memory recurrent neural network architectures for large scale acoustic modeling
Sak, H.; Senior, A. W.; and Beaufays, F. 2014 · 2014
Earlier work this paper cites.
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; and Mordatch, I. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Counterfactual Multi-Agent Policy Gradients
Foerster, J. N.; Farquhar, G.; Afouras, T.; Nardelli, N.; and Whiteson, S. 2018 · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
Haarnoja, T.; Zhou, A.; Hartikainen, K.; Tucker, G.; Ha, S.; Tan, J.; Kumar, V.; Zhu, H.; Gupta, A.; Abbeel, P.; et al. 2018 · 2018
Cited alongside, same era.
Self-attention with relative position representations
Shaw, P.; Uszkoreit, J.; and Vaswani, A. 2018 · 2018
Cited alongside, same era.
RUDDER: Return Decomposition for Delayed Rewards
Arjona-Medina, J. A.; Gillhofer, M.; Widrich, M.; Unterthiner, T.; Brandstetter, J.; and Hochreiter, S. 2019 · 2019
Cited alongside, same era.
LIIR: learning individual intrinsic reward in multi-agent reinforcement learning
Du, Y.; Han, L.; Fang, M.; Dai, T.; Liu, J.; and Tao, D. 2019 · 2019
Cited alongside, same era.
Shapley Counterfactual Credits for Multi-Agent Reinforcement Learning
Li, J.; Kuang, K.; Wang, B.; Liu, F.; Chen, L.; Wu, F.; and Xiao, J. 2021 · 2021
Later among the works it cites.
Multi-agent reinforcement learning for redundant robot control in task-space
Perrusquía, A.; Yu, W.; and Li, X. 2021 · 2021
Later among the works it cites.
DOP: Off-Policy Multi-Agent Decomposed Policy Gradients
Wang, Y.; Han, B.; Wang, T.; Dong, H.; and Zhang, C. 2021 · 2021
Later among the works it cites.
Off-Policy Reinforcement Learning with Delayed Rewards
Han, B.; Ren, Z.; Wu, Z.; Zhou, Y.; and Peng, J. 2022 · 2022
Later among the works it cites.
Align-RUDDER: Learning From Few Demonstrations by Reward Redistribution
Patil, V. P.; Hofmarcher, M.; Dinu, M.; Dorfer, M.; Blies, P. M.; Brandstetter, J.; Arjona-Medina, J. A.; and Hochreiter, S. 2022 · 2022
Later among the works it cites.
Learning Long-Term Reward Redistribution via Randomized Return Decomposition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data shapley: Equitable valuation of data for machine learning
Ghorbani, A.; and Zou, J. 2019 · 2019
Cited alongside, same era.
QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning
Son, K.; Kim, D.; Kang, W. J.; Hostallero, D.; and Yi, Y. 2019 · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O.; Babuschkin, I.; Czarnecki, W. M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D. H.; Powell, R.; Ewalds, T.; Georgiev, P.; et al. 2019 · 2019
Cited alongside, same era.
Learning Guidance Rewards with Trajectory-space Smoothing
Gangwani, T.; Zhou, Y.; and Peng, J. 2020 · 2020
Cited alongside, same era.
Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning
Rashid, T.; Farquhar, G.; Peng, B.; and Whiteson, S. 2020 · 2020
Cited alongside, same era.
The many Shapley values for model explanation
Sundararajan, M.; and Najmi, A. 2020 · 2020
Cited alongside, same era.
Shapley Q-Value: A Local Reward Approach to Solve Global Reward Games
Wang, J.; Zhang, Y.; Kim, T.; and Gu, Y. 2020a · 2020
Cited alongside, same era.
Ren, Z.; Guo, R.; Zhou, Y.; and Peng, J. 2022 · 2022
Later among the works it cites.
Agent-time attention for sparse rewards multi-agent reinforcement learning
She, J.; Gupta, J. K.; and Kochenderfer, M. J. 2022 · 2022
Later among the works it cites.
Agent-Temporal Attention for Reward Redistribution in Episodic Multi-Agent Reinforcement Learning
Xiao, B.; Ramasubramanian, B.; and Poovendran, R. 2022 · 2022
Later among the works it cites.
A Review of Cooperation in Multi-agent Learning
Du, Y.; Leibo, J. Z.; Islam, U.; Willis, R.; and Sunehag, P. 2023 · 2023
Closest in time.
Interpretable Reward Redistribution in Reinforcement Learning: A Causal Approach
Zhang, Y.; Du, Y.; Huang, B.; Wang, Z.; Wang, J.; Fang, M.; and Pechenizkiy, M. 2023 · 2023
Closest in time.
Probably approximate Shapley fairness with applications in machine learning
Zhou, Z.; Xu, X.; Sim, R. H. L.; Foo, C. S.; and Low, B. K. H. 2023 · 2023
Closest in time.
Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward
Sunehag, P.; Lever, G.; Gruslys, A.; Czarnecki, W. M.; Zambaldi, V. F.; Jaderberg, M.; Lanctot, M.; Sonnerat, N.; Leibo, J. Z.; Tuyls, K.; and Graepel, T. 2018 · 2087
Closest in time.