Fetching the paper…
Reading the bibliography…
This paper considers multi-agent reinforcement learning (MARL) tasks where agents receive a shared global reward at the end of an episode.
Neural networks and the bias/variance dilemma
Stuart Geman, Elie Bienenstock, and René Doursat. 1992 · 1992
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping. In International Conference on Machine Learning . 278–287
Andrew Y Ng, Daishi Harada, and Stuart Russell. 1999 · 1999
Earlier work this paper cites.
A unified bias-variance decomposition for zero-one and squared loss
Pedro Domingos. 2000 · 2000
Earlier work this paper cites.
The Elements of Statistical Learning
Jerome Friedman, Trevor Hastie, and Robert Tibshirani. 2001 · 2001
Earlier work this paper cites.
QUICR-learning for multi-agent coordination. In Proceedings of the National Conference on Artificial Intelligence , Vol. 21. 1438
Adrian K Agogino and Kagan Tumer. 2006 · 2006
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter. 2008 · 2008
Earlier work this paper cites.
Theoretical considerations of potential-based reward shaping for multi-agent systems. In International Conference on Autonomous Agents and Multi-Agent Systems . 225–232
Sam Devlin and Daniel Kudenko. 2011 · 2011
Earlier work this paper cites.
An empirical study of potential-based reward shaping and advice in complex, multi-agent systems
Sam Devlin, Daniel Kudenko, and Marek Grześ. 2011 · 2011
Earlier work this paper cites.
Policy invariance under reward transformations for general-sum stochastic games
Xiaosong Lu, Howard M Schwartz, and Sidney Nascimento Givigi. 2011 · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning. In International Conference on Artificial Intelligence and Statistics . 627–635
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell. 2011 · 2011
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling. 2013 · 2013
Earlier work this paper cites.
Potential-based difference rewards for multiagent reinforcement learning. In International Conference on Autonomous Agents and Multi-Agent Systems . 165–172
Sam Devlin, Logan Yliniemi, Daniel Kudenko, and Kagan Tumer. 2014 · 2014
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
A concise introduction to decentralized POMDPs . Vol. 1
Frans A Oliehoek, Christopher Amato, et al · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis. 2016 · 2016
Earlier work this paper cites.
Cooperative multi-agent control using deep reinforcement learning. In International Conference on Autonomous Agents and Multiagent Systems . 66–83
Jayesh K Gupta, Maxim Egorov, and Mykel Kochenderfer. 2017 · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments. In Neural Information Processing Systems . 6379–6390
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Cited alongside, same era.
Molecular de-novo design through deep reinforcement learning
Marcus Olivecrona, Thomas Blaschke, Ola Engkvist, and Hongming Chen. 2017 · 2017
Cited alongside, same era.
Deep reinforcement learning framework for autonomous driving
Ahmad EL Sallab, Mohammed Abdou, Etienne Perot, and Senthil Yogamani. 2017 · 2017
Cited alongside, same era.
Multiagent cooperation and competition with deep reinforcement learning
Ardi Tampuu, Tambet Matiisen, Dorian Kodelja, Ilya Kuzovkin, Kristjan Korjus, Juhan Aru, Jaan Aru, and Raul Vicente. 2017 · 2017
Cited alongside, same era.
Attention is all you need. In Neural Information Processing Systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Graph Convolutional Reinforcement Learning. In International Conference on Learning Representations
Jiechuan Jiang, Chen Dun, Tiejun Huang, and Zongqing Lu. 2019 · 2019
Later among the works it cites.
HG-DAgger: Interactive Imitation Learning with Human Experts. In International Conference on Robotics and Automation . IEEE, 8077–8083
Michael Kelly, Chelsea Sidrane, Katherine Driggs-Campbell, and Mykel J Kochenderfer. 2019 · 2019
Later among the works it cites.
Sequence Modeling of Temporal Credit Assignment for Episodic Reinforcement Learning
Yang Liu, Yunan Luo, Yuanyi Zhong, Xi Chen, Qiang Liu, and Jian Peng. 2019 · 2019
Later among the works it cites.
MAVEN: Multi-agent variational exploration
Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, and Shimon Whiteson. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep Sets. In Neural Information Processing Systems
Manzil Zaheer, Satwik Kottur, Siamak Ravanbhakhsh, Barnabás Póczos, Ruslan Salakhutdinov, and Alexander J Smola. 2017 · 2017
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients. In AAAI Conference on Artificial Intelligence . 2974–2982
Jakob N Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. 2018 · 2018
Cited alongside, same era.
Learning attentional communication for multi-agent cooperation. In Neural Information Processing Systems
Jiechuan Jiang and Zongqing Lu. 2018 · 2018
Cited alongside, same era.
A modern take on the bias-variance tradeoff in neural networks
Brady Neal, Sarthak Mittal, Aristide Baratin, Vinayak Tantia, Matthew Scicluna, Simon Lacoste-Julien, and Ioannis Mitliagkas. 2018 · 2018
Cited alongside, same era.
QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning. In International Conference on Machine Learning . 4295–4304
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. 2018 · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Cited alongside, same era.
Hangyu Mao, Zhengchao Zhang, Zhen Xiao, and Zhibo Gong. 2019 · 2019
Later among the works it cites.
The StarCraft Multi-Agent Challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philiph H. S. Torr, Jakob Foerster, and Shimon Whiteson. 2019 · 2019
Later among the works it cites.
QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning. In International Conference on Machine Learning . 5887–5896
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi. 2019 · 2019
Later among the works it cites.
High-dimensional statistics: A non-asymptotic viewpoint
Martin J Wainwright. 2019 · 2019
Later among the works it cites.
Learning Guidance Rewards with Trajectory-space Smoothing. In Neural Information Processing Systems
Tanmay Gangwani, Yuan Zhou, and Jian Peng. 2020 · 2020
Later among the works it cites.
MA-TREX: Multi-agent Trajectory-Ranked Reward Extrapolation via Inverse Reinforcement Learning. In International Conference on Knowledge Science, Engineering and Management . Springer, 3–14
Sili Huang, Bo Yang, Hechang Chen, Haiyin Piao, Zhixiao Sun, and Yi Chang. 2020 · 2020
Later among the works it cites.
Multi-agent actor-critic with hierarchical graph attention network. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 7236–7243
Heechang Ryu, Hayong Shin, and Jinkyoo Park. 2020 · 2020
Later among the works it cites.
Shapley Q-value: A local reward approach to solve global reward games. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 7285–7292
Jianhong Wang, Yuan Zhang, Tae-Kyun Kim, and Yunjie Gu. 2020 · 2020
Later among the works it cites.
Q-value Path Decomposition for Deep Multiagent Reinforcement Learning. In International Conference on Machine Learning
Yaodong Yang, Jianye Hao, Guangyong Chen, Hongyao Tang, Yingfeng Chen, Yujing Hu, Changjie Fan, and Zhongyu Wei. 2020 · 2020
Later among the works it cites.
Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning. In Neural Information Processing Systems
Meng Zhou, Ziyu Liu, Pengwei Sui, Yixuan Li, and Yuk Ying Chung. 2020 · 2020
Later among the works it cites.
Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward.. In International Conference on Autonomous Agents and Multiagent Systems . 2085–2087
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinícius Flores Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2087
Closest in time.