Fetching the paper…
Reading the bibliography…
Modelling and exploiting teammates' policies in cooperative multi-agent systems have long been an interest and also a big challenge for the reinforcement learning (RL) community.
Approximation by superpositions of a sigmoidal function
George Cybenko. 1989 · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White. 1989 · 1989
Earlier work this paper cites.
Opponent modeling in poker
Darse Billings, Denis Papp, Jonathan Schaeffer, and Duane Szafron. 1998 · 1998
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton and Andrew G Barto. 1998 · 1998
Earlier work this paper cites.
TCP-like congestion control for layered multicast data transfer. In IEEE infocom
Lorenzo Vicisano, Luigi Rizzo, and Jon Crowcroft. 1998 · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation. In Advances in neural information processing systems
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour. 2000 · 2000
Earlier work this paper cites.
A multi-agent, policy-gradient approach to network routing. In In: Proc. of the 18th Int. Conf. on Machine Learning
Nigel Tao, Jonathan Baxter, and Lex Weaver. 2001 · 2001
Earlier work this paper cites.
The complexity of decentralized control of Markov decision processes
Daniel S Bernstein, Robert Givan, Neil Immerman, and Shlomo Zilberstein. 2002 · 2002
Earlier work this paper cites.
On actor-critic algorithms
Vijay R Konda and John N Tsitsiklis. 2003 · 2003
Earlier work this paper cites.
Best-response multiagent learning in non-stationary environments. In Proceedings of the Third International Joint Conference on Autonomous Agents and Multiagent Systems-Volume 2
Michael Weinberg and Jeffrey S Rosenschein. 2004 · 2004
Earlier work this paper cites.
Walking the tightrope: Responsive yet stable traffic engineering. In ACM SIGCOMM Computer Communication Review
Srikanth Kandula, Dina Katabi, Bruce Davie, and Anna Charny. 2005 · 2005
Earlier work this paper cites.
Reasoning about joint beliefs for execution-time communication decisions. In Proceedings of the fourth international joint conference on Autonomous agents and multiagent systems
Maayan Roth, Reid Simmons, and Manuela Veloso. 2005 · 2005
Earlier work this paper cites.
Reaching pareto-optimality in prisoner¡¯s dilemma using conditional joint action learning
Dipyaman Banerjee and Sandip Sen. 2007 · 2007
Earlier work this paper cites.
A multiagent approach to autonomous intersection management
Kurt Dresner and Peter Stone. 2008 · 2008
Earlier work this paper cites.
Optimal and approximate Q-value functions for decentralized POMDPs
Frans A Oliehoek, Matthijs TJ Spaan, and Nikos Vlassis. 2008 · 2008
Earlier work this paper cites.
Game theory-based opponent modeling in large imperfect-information games. In The 10th International Conference on Autonomous Agents and Multiagent Systems-Volume 2
Sam Ganzfried and Tuomas Sandholm. 2011 · 2011
Cited alongside, same era.
A survey of actor-critic reinforcement learning: Standard and natural policy gradients
Ivo Grondman, Lucian Busoniu, Gabriel AD Lopes, and Robert Babuska. 2012 · 2012
Cited alongside, same era.
Coordinating multi-agent reinforcement learning with limited communication. In Proceedings of the 2013 international conference on Autonomous agents and multi-agent systems
Chongjie Zhang and Victor Lesser. 2013 · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms. In ICML
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014 · 2014
Cited alongside, same era.
WCMP: Weighted cost multipathing for improved fairness in data centers. In Proceedings of the Ninth European Conference on Computer Systems
Fully decentralized policies for multi-agent systems: An information theoretic approach. In Advances in Neural Information Processing Systems
Roel Dobbe, David Fridovich-Keil, and Claire Tomlin. 2017 · 2017
Later among the works it cites.
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. 2017 · 2017
Later among the works it cites.
Cooperative multi-agent control using deep reinforcement learning. In International Conference on Autonomous Agents and Multiagent Systems
Jayesh K Gupta, Maxim Egorov, and Mykel Kochenderfer. 2017 · 2017
Later among the works it cites.
A survey of learning in multiagent environments: Dealing with non-stationarity
Pablo Hernandez-Leal, Michael Kaisers, Tim Baarslag, and Enrique Munoz de Cote. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Junlan Zhou, Malveeka Tewari, Min Zhu, Abdul Kabbani, Leon Poutievski, Arjun Singh, and Amin Vahdat. 2014 · 2014
Cited alongside, same era.
Efficient traffic splitting on sdn switches. In Proceedings of CoNEXT
Nanxi Kang, Monia Ghobadi, John Reumann, Alexander Shraer, and Jennifer Rexford. 2015 · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015 · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Universal value function approximators. In International Conference on Machine Learning
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver. 2015 · 2015
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Van Hasselt, Marc Lanctot, and Nando De Freitas. 2015 · 2015
Cited alongside, same era.
Opponent modeling in deep reinforcement learning. In International Conference on Machine Learning
He He, Jordan Boyd-Graber, Kevin Kwok, and Hal Daumé III. 2016 · 2016
Cited alongside, same era.
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Later among the works it cites.
Hangyu Mao, Zhibo Gong, Yan Ni, and Zhen Xiao. 2017 · 2017
Later among the works it cites.
Value-decomposition networks for cooperative multi-agent learning
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2017
Later among the works it cites.
Autonomous agents modelling other agents: A comprehensive survey and open problems
Stefano V Albrecht and Peter Stone. 2018 · 2018
Closest in time.
Learning with opponent-learning awareness. In Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems
Jakob Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch. 2018 · 2018
Closest in time.
A Deep Policy Inference Q-Network for Multi-Agent Systems. In Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems
Zhang-Wei Hong, Shih-Yang Su, Tzu-Yun Shann, Yi-Hsiang Chang, and Chun-Yi Lee. 2018 · 2018
Closest in time.
Modeling Others using Oneself in Multi-Agent Reinforcement Learning
Roberta Raileanu, Emily Denton, Arthur Szlam, and Rob Fergus. 2018 · 2018
Closest in time.
QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder de Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. 2018 · 2018
Closest in time.
Mean Field Multi-Agent Reinforcement Learning
Yaodong Yang, Rui Luo, Minne Li, Ming Zhou, Weinan Zhang, and Jun Wang. 2018 · 2018
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention. In International conference on machine learning
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.