Fetching the paper…
Reading the bibliography…
This paper surveys the field of deep multiagent reinforcement learning.
Mnih V, Badia AP, Mirza M, Graves A, Lillicrap T, Harley T, Silver D, Kavukcuoglu K (2016) Asynchronous methods for deep reinforcement learning. In: Balcan MF, Weinberger KQ (eds) Proceedings of The 33rd International Conference on Machine Learning, PMLR, pp 1928–1937
1937
Earlier work this paper cites.
Brown GW (1951) Iterative solution of games by fictitious play. Activity Analysis of Production and Allocation 13(1):374–376
1951
Earlier work this paper cites.
Kuhn HW, Tucker AW (1953) Contributions to the theory of games, vol 2. Princeton University Press
1953
Earlier work this paper cites.
Shapley LS (1953) Stochastic games. Proceedings of the National Academy of Sciences 39(10):1095–1100
1953
Earlier work this paper cites.
Bellman R (1957) A markovian decision process. Journal of Mathematics and Mechanics pp 679–684
1957
Earlier work this paper cites.
Simon HA (1957) Models of man, social and rational: Mathematical essays on rational human behavior in a social setting. Wiley and Sons
1957
Earlier work this paper cites.
Srivastava N, Hinton G, Krizhevsky A, Sutskever I, Salakhutdinov R (2014) Dropout: a simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research 15(1):1929–1958
1958
Earlier work this paper cites.
Minsky M (1961) Steps toward artificial intelligence. Proceedings of the IRE 49(1):8–30, DOI 10.1109/JRPROC.1961.287775
1961
Earlier work this paper cites.
Åström KJ (1965) Optimal control of markov decision processes with incomplete state estimation. Journal of Mathematical Analysis and Applications 10:174–205
1965
Earlier work this paper cites.
Stanley HE (1971) Phase transitions and critical phenomena. Clarendon, Oxford
1971
Earlier work this paper cites.
Premack D, Woodruff G (1978) Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences 1(4):515–526
1978
Earlier work this paper cites.
Axelrod R, Hamilton WD (1981) The evolution of cooperation. Science 211(4489):1390–1396
1981
Earlier work this paper cites.
Dovidio JF (1984) Helping behavior and altruism: An empirical and conceptual overview. Advances in Experimental Social Psychology 17:361–427
1984
Earlier work this paper cites.
Simon HA (1990) Bounded rationality. Springer
1990
Earlier work this paper cites.
Williams RJ (1992) Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning 8(3-4):229–256
1992
Earlier work this paper cites.
Bäck T, Schwefel HP (1993) An overview of evolutionary algorithms for parameter optimization. Evolutionary Computation 1(1):1–23
1993
Earlier work this paper cites.
Tan M (1993) Multi-agent reinforcement learning: Independent vs. cooperative agents. In: Proceedings of the tenth International Conference on Machine Learning, pp 330–337
1993
Earlier work this paper cites.
Littman ML (1994) Markov games as a framework for multi-agent reinforcement learning. In: 11th International Conference on Machine Learning, Elsevier, pp 157–163
1994
Earlier work this paper cites.
Gigerenzer G, Goldstein DG (1996) Reasoning the fast and frugal way: models of bounded rationality. Psychological Review 103(4):650
1996
Earlier work this paper cites.
Sutton RS, Barto AG, et al. (1998) Introduction to reinforcement learning, vol 135. MIT press Cambridge
1998
Earlier work this paper cites.
Fehr E, Schmidt KM (1999) A theory of fairness, competition, and cooperation. The Quarterly Journal of Economics 114(3):817–868
1999
Earlier work this paper cites.
Moriarty DE, Schultz AC, Grefenstette JJ (1999) Evolutionary algorithms for reinforcement learning. Journal of Artificial Intelligence Research 11:241–276
1999
Earlier work this paper cites.
Ng AY, Harada D, Russell S (1999) Policy invariance under reward transformations: Theory and application to reward shaping. In: ICML, vol 99, pp 278–287
1999
Earlier work this paper cites.
Sutton RS, McAllester D, Singh S, Mansour Y (1999) Policy gradient methods for reinforcement learning with function approximation. Advances in neural information processing systems 12
1999
Earlier work this paper cites.
Bowling M, Veloso M (2001) Rational and convergent learning in stochastic games. In: International Joint Conference on Artificial Intelligence, Citeseer, vol 17, pp 1021–1026
2001
Earlier work this paper cites.
Amato C, Oliehoek F (2015) Scalable planning and learning for multiagent pomdps. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 29, pp 1995–2002
2002
Earlier work this paper cites.
Bernstein DS, Givan R, Immerman N, Zilberstein S (2002) The complexity of decentralized control of markov decision processes. Mathematics of Operations Research 27(4):819–840
2002
Earlier work this paper cites.
Bowling M, Veloso M (2002) Multiagent learning using a variable learning rate. Artificial Intelligence 136(2):215–250
2002
Earlier work this paper cites.
Gilovich T, Griffin D, Kahneman D (2002) Heuristics and biases: The psychology of intuitive judgment. Cambridge University Press
2002
Earlier work this paper cites.
Colman AM (2003) Cooperation, psychological game theory, and limitations of rationality in social interaction. Behavioral and Brain Sciences 26:139–198
2003
Earlier work this paper cites.
Kakade SM (2003) On the sample complexity of reinforcement learning. University of London, University College London (United Kingdom)
2003
Earlier work this paper cites.
Konda VR, Tsitsiklis JN (2003) Actor-critic algorithms. Journal on Control and Optimization 42(4):1143–1166
2003
Earlier work this paper cites.
Greensmith E, Bartlett PL, Baxter J (2004) Variance reduction techniques for gradient estimates in reinforcement learning. Journal of Machine Learning Research 5(9)
2004
Earlier work this paper cites.
Hansen EA, Bernstein DS, Zilberstein S (2004) Dynamic programming for partially observable stochastic games. In: American Association for Artificial Intelligence, vol 4, pp 709–715
2004
Earlier work this paper cites.
Frith C, Frith U (2005) Theory of mind. Current Biology 15(17):644–645
2005
Earlier work this paper cites.
Markovitch S, Reger R (2005) Learning and exploiting relative weaknesses of opponent agents. Autonomous Agents and Multi-Agent Systems 10(2):103–130
2005
Earlier work this paper cites.
Nevmyvaka Y, Feng Y, Kearns M (2006) Reinforcement learning for optimized trade execution. In: Proceedings of the 23rd international conference on Machine learning, pp 673–680
2006
Earlier work this paper cites.
Busoniu L, Babuska R, De Schutter B (2008) A comprehensive survey of multiagent reinforcement learning. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) 38(2):156–172
2008
Earlier work this paper cites.
Lehman J, Stanley KO (2008) Exploiting open-endedness to solve problems through the search for novelty. In: Artificial Life XI, Citeseer, pp 329–336
2008
Earlier work this paper cites.
Peters J, Schaal S (2008) Natural actor-critic. Neurocomputing 71(7-9):1180–1190
2008
Earlier work this paper cites.
Shoham Y, Leyton-Brown K (2008) Multiagent systems: Algorithmic, game-theoretic, and logical foundations. Cambridge University Press
2008
Earlier work this paper cites.
Kumar A, Zilberstein S (2009) Dynamic programming approximations for partially observable stochastic games. In: Proceedings of the Twenty-Second International FLAIRS Conference, pp 547–552
2009
Earlier work this paper cites.
Taylor ME, Stone P (2009) Transfer learning for reinforcement learning domains: A survey. Journal of Machine Learning Research 10(7)
2009
Earlier work this paper cites.
Marewski JN, Gaissmaier W, Gigerenzer G (2010) Good judgments do not require complex cognition. Cognitive Processing 11(2):103–121
2010
Earlier work this paper cites.
Devlin S, Kudenko D (2011) Theoretical considerations of potential-based reward shaping for multi-agent systems. In: The 10th International Conference on Autonomous Agents and Multiagent Systems, ACM, pp 225–232
2011
Earlier work this paper cites.
Devlin S, Kudenko D, Grześ M (2011) An empirical study of potential-based reward shaping and advice in complex, multi-agent systems. Advances in Complex Systems 14(02):251–278
2011
Earlier work this paper cites.
Devlin SM, Kudenko D (2012) Dynamic potential-based reward shaping. In: Proceedings of the 11th International Conference on Autonomous Agents and Multiagent Systems, IFAAMAS, pp 433–440
2012
Earlier work this paper cites.
Grondman I, Busoniu L, Lopes GA, Babuska R (2012) A survey of actor-critic reinforcement learning: Standard and natural policy gradients. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) 42(6):1291–1307
2012
Earlier work this paper cites.
Nitschke GS, Eiben A, Schut MC (2012) Evolving team behaviors with specialization. Genetic Programming and Evolvable Machines 13(4):493–536
2012
Earlier work this paper cites.
Proper S, Tumer K (2012) Modeling difference rewards for multiagent learning. In: Proceedings of the 11th International Conference on Autonomous Agents and Multiagent Systems), Conitzer, Winikoff, Padgham, pp 1397–1398
2012
Earlier work this paper cites.
Van Otterlo M, Wiering M (2012) Reinforcement learning and markov decision processes. In: Reinforcement learning, Springer, pp 3–42
2012
Earlier work this paper cites.
Johanson M, Burch N, Valenzano R, Bowling M (2013) Evaluating state-space abstractions in extensive-form games. In: Proceedings of the 2013 International Conference on Autonomous Agents and Multi-agent Systems, pp 271–278
2013
Earlier work this paper cites.
Mnih V, Kavukcuoglu K, Silver D, Graves A, Antonoglou I, Wierstra D, Riedmiller M (2013) Playing atari with deep reinforcement learning. arXiv preprint arXiv:13125602
2013
Earlier work this paper cites.
Van Der Ree M, Wiering M (2013) Reinforcement learning in the game of othello: Learning against a fixed opponent and learning from self-play. In: 2013 IEEE Symposium on Adaptive Dynamic Programming And Reinforcement Learning (ADPRL), IEEE, pp 108–115
2013
Earlier work this paper cites.
Devlin S, Yliniemi L, Kudenko D, Tumer K (2014) Potential-based difference rewards for multiagent reinforcement learning. In: Proceedings of the 2014 International Conference on Autonomous Agents and Multi-agent Systems, pp 165–172
2014
Earlier work this paper cites.
Gomes J, Mariano P, Christensen AL (2014) Avoiding convergence in cooperative coevolution with novelty search. In: Proceedings of the 2014 International Conference on Autonomous Agents and Multi-agent Systems, pp 1149–1156
2014
Earlier work this paper cites.
Silver D, Lever G, Heess N, Degris T, Wierstra D, Riedmiller M (2014) Deterministic policy gradient algorithms. In: International conference on machine learning, PMLR, pp 387–395
2014
Earlier work this paper cites.
Yliniemi L, Tumer K (2014) Multi-objective multiagent credit assignment through difference rewards in reinforcement learning. In: Asia-Pacific Conference on Simulated Evolution and Learning, Springer, pp 407–418
2014
Earlier work this paper cites.
Bloembergen D, Tuyls K, Hennes D, Kaisers M (2015) Evolutionary dynamics of multi-agent learning: A survey. Journal of Artificial Intelligence Research 53:659–697
2015
Earlier work this paper cites.
Bowling M, Burch N, Johanson M, Tammelin O (2015) Heads-up limit hold’em poker is solved. Science 347(6218):145–149
2015
Earlier work this paper cites.
Hausknecht M, Stone P (2015) Deep recurrent q-learning for partially observable mdps. In: 2015 aaai fall symposium series
2015
Earlier work this paper cites.
Heinrich J, Lanctot M, Silver D (2015) Fictitious self-play in extensive-form games. In: International Conference on Machine Learning, PMLR, pp 805–813
2015
Earlier work this paper cites.
LeCun Y, Bengio Y, Hinton G (2015) Deep learning. Nature 521(7553):436–444
2015
Earlier work this paper cites.
Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG, Graves A, Riedmiller M, Fidjeland AK, Ostrovski G, et al. (2015) Human-level control through deep reinforcement learning. Nature 518(7540):529–533
2015
Earlier work this paper cites.
Schulman J, Levine S, Abbeel P, Jordan M, Moritz P (2015) Trust region policy optimization. In: International Conference on Machine Learning, PMLR, pp 1889–1897
2015
Cited alongside, same era.
Wen Z, O’Neill D, Maei H (2015) Optimal demand response using device-based reinforcement learning. IEEE Transactions on Smart Grid 6(5):2312–2324
2015
Cited alongside, same era.
Amir O, Kamar E, Kolobov A, Grosz B (2016) Interactive teaching strategies for agent training. In: Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence 2016, URL https://www.microsoft.com/en-us/research/publication/interactive-teaching-strategies-agent-training/
2016
Cited alongside, same era.
Colin TR, Belpaeme T, Cangelosi A, Hemion N (2016) Hierarchical reinforcement learning as creative problem solving. Robotics and Autonomous Systems 86:196–206
2016
Cited alongside, same era.
Bao W, Liu Xy (2019) Multi-agent deep reinforcement learning for liquidation strategy analysis. arXiv preprint arXiv:190611046
2019
Later among the works it cites.
Berner C, Brockman G, Chan B, Cheung V, Debiak P, Dennison C, Farhi D, Fischer Q, Hashme S, Hesse C, Józefowicz R, Gray S, Olsson C, Pachocki JW, Petrov M, de Oliveira Pinto HP, Raiman J, Salimans T, Schlatter J, Schneider J, Sidor S, Sutskever I, Tang J, Wolski F, Zhang S (2019) Dota 2 with large scale deep reinforcement learning. arXiv preprint arXiv:191206680
2019
Later among the works it cites.
Brown N, Sandholm T (2019) Superhuman ai for multiplayer poker. Science 365(6456):885–890
2019
Later among the works it cites.
Da Silva FL, Costa AHR (2019) A survey on transfer learning for multiagent reinforcement learning systems. Journal of Artificial Intelligence Research 64:645–703
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Foerster J, Assael IA, De Freitas N, Whiteson S (2016) Learning to communicate with deep multi-agent reinforcement learning. Advances in Neural Information Processing Systems 29:2137–2145
2016
Cited alongside, same era.
Hausknecht M, Stone P (2016) Grounded semantic networks for learning shared communication protocols. In: International Conference on Machine Learning (Workshop)
2016
Cited alongside, same era.
He H, Boyd-Graber J, Kwok K, Daumé III H (2016) Opponent modeling in deep reinforcement learning. In: International Conference on Machine Learning, Proceedings of Machine Learning Research, pp 1804–1813
2016
Cited alongside, same era.
Heinrich J, Silver D (2016) Deep reinforcement learning from self-play in imperfect-information games. arXiv preprint arXiv:160301121
2016
Cited alongside, same era.
Hernandez-Leal P, Rosman B, Taylor ME, Sucar LE, Munoz de Cote E (2016) A bayesian approach for learning and tracking switching, non-stationary opponents. In: Proceedings of the 2016 International Conference on Autonomous Agents & Multiagent Systems, pp 1315–1316
2016
Cited alongside, same era.
Holmesparker C, Agogino AK, Tumer K (2016) Combining reward shaping and hierarchies for scaling to large multiagent systems. The Knowledge Engineering Review 31(1):3–18
2016
Cited alongside, same era.
Kraemer L, Banerjee B (2016) Multi-agent reinforcement learning as a rehearsal for decentralized planning. Neurocomputing 190:82–94
2016
Cited alongside, same era.
Kurek M, Jaśkowski W (2016) Heterogeneous team deep q-learning in low-dimensional multi-agent environments. In: 2016 IEEE Conference on Computational Intelligence and Games (CIG), IEEE, pp 1–8
2016
Cited alongside, same era.
Dankwa S, Zheng W (2019) Twin-delayed ddpg: A deep reinforcement learning technique to model a continuous movement of an intelligent robot agent. In: Proceedings of the 3rd International Conference on Vision, Image and Signal Processing, pp 1–5
2019
Later among the works it cites.
Drugan MM (2019) Reinforcement learning versus evolutionary computation: A survey on hybrid algorithms. Swarm and Evolutionary Computation 44:228–246
2019
Later among the works it cites.
Du Y, Han L, Fang M, Liu J, Dai T, Tao D (2019) Liir: Learning individual intrinsic reward in multi-agent reinforcement learning. Advances in Neural Information Processing Systems 32:4403–4414
2019
Later among the works it cites.
Eccles T, Hughes E, Kramár J, Wheelwright S, Leibo JZ (2019) Learning reciprocity in complex sequential social dilemmas. arXiv preprint arXiv:190308082
2019
Later among the works it cites.
Graesser L, Keng WL (2019) Foundations of deep reinforcement learning: theory and practice in Python. Addison-Wesley Professional
2019
Later among the works it cites.
Hernandez-Leal P, Kartal B, Taylor ME (2019) A survey and critique of multiagent deep reinforcement learning. Autonomous Agents and Multi-Agent Systems 33(6):750–797
2019
Later among the works it cites.
Ilhan E, Gow J, Perez-Liebana D (2019) Teaching on a budget in multi-agent deep reinforcement learning. In: 2019 IEEE Conference on Games (CoG), IEEE, pp 1–8
2019
Later among the works it cites.
Iqbal S, Sha F (2019) Actor-attention-critic for multi-agent reinforcement learning. In: International Conference on Machine Learning, PMLR, pp 2961–2970
2019
Later among the works it cites.
Jaderberg M, Czarnecki WM, Dunning I, Marris L, Lever G, Castaneda AG, Beattie C, Rabinowitz NC, Morcos AS, Ruderman A, et al. (2019) Human-level performance in 3d multiplayer games with population-based reinforcement learning. Science 364(6443):859–865
2019
Later among the works it cites.
Jaques N, Lazaridou A, Hughes E, Gulcehre C, Ortega P, Strouse D, Leibo JZ, De Freitas N (2019) Social influence as intrinsic motivation for multi-agent deep reinforcement learning. In: International Conference on Machine Learning, PMLR, pp 3040–3049
2019
Later among the works it cites.
Kim W, Cho M, Sung Y (2019) Message-dropout: An efficient training method for multi-agent deep reinforcement learning. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 33, pp 6079–6086, DOI https://doi.org/10.1609/aaai.v33i01.33016079
2019
Later among the works it cites.
Li S, Wu Y, Cui X, Dong H, Fang F, Russell S (2019) Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 33, pp 4213–4220
2019
Later among the works it cites.
Liu S, Lever G, Merel J, Tunyasuvunakool S, Heess N, Graepel T (2019) Emergent coordination through competition. arXiv preprint arXiv:190207151
2019
Later among the works it cites.
Lowe R, Foerster J, Boureau YL, Pineau J, Dauphin Y (2019) On the pitfalls of measuring emergent communication. In: Proceedings of the 18th International Conference on Autonomous Agents and Multiagent Systems, International Foundation for Autonomous Agents and Multi Agent Systems, Richland, SC, AAMAS ’19, p 693–701
2019
Later among the works it cites.
Mahajan A, Rashid T, Samvelyan M, Whiteson S (2019) Maven: Multi-agent variational exploration. Advances in Neural Information Processing Systems 32
2019
Later among the works it cites.
Omidshafiei S, Kim DK, Liu M, Tesauro G, Riemer M, Amato C, Campbell M, How JP (2019) Learning to teach in cooperative multiagent reinforcement learning. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 33, pp 6128–6136
2019
Later among the works it cites.
Papoudakis G, Christianos F, Rahman A, Albrecht SV (2019) Dealing with non-stationarity in multi-agent deep reinforcement learning. arXiv preprint arXiv:190604737
2019
Later among the works it cites.
Prasad A, Dusparic I (2019) Multi-agent deep reinforcement learning for zero energy communities. In: 2019 IEEE PES innovative smart grid technologies Europe (ISGT-Europe), IEEE, pp 1–5
2019
Later among the works it cites.
Son K, Kim D, Kang WJ, Hostallero DE, Yi Y (2019) Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning. In: International Conference on Machine Learning, PMLR, pp 5887–5896
2019
Later among the works it cites.
Vinyals O, Babuschkin I, Czarnecki WM, Mathieu M, Dudzik A, Chung J, Choi DH, Powell R, Ewalds T, Georgiev P, et al. (2019) Grandmaster level in starcraft ii using multi-agent reinforcement learning. Nature 575(7782):350–354
2019
Later among the works it cites.
Wang RE, Everett M, How JP (2019) R-maddpg for partially observable environments and limited communication. In: International Conference on Machine Learning 2019 Workshop RL4RealLife
2019
Later among the works it cites.
Wen Y, Yang Y, Luo R, Wang J, Pan W (2019) Probabilistic recursive reasoning for multi-agent reinforcement learning. In: 7th International Conference on Learning Representations, ICLR 2019
2019
Later among the works it cites.
Schroeder de Witt C, Foerster J, Farquhar G, Torr P, Boehmer W, Whiteson S (2019) Multi-agent common knowledge reinforcement learning. Advances in Neural Information Processing Systems 32:9927–9939
2019
Later among the works it cites.
Burden J (2020) Automating abstraction for potential-based reward shaping. PhD thesis, University of York
2020
Later among the works it cites.
Chu T, Wang J, Codecà L, Li Z (2020) Multi-agent deep reinforcement learning for large-scale traffic signal control. IEEE Transactions on Intelligent Transportation Systems 21(3):1086–1095
2020
Later among the works it cites.
Dai Z, Chen Y, Low BKH, Jaillet P, Ho TH (2020) R2-b2: Recursive reasoning-based bayesian optimization for no-regret learning in games. In: Proceedings of the 37th International Conference on Machine Learning, PMLR, pp 2291–2301
2020
Later among the works it cites.
Ding Z, Dong H (2020) Challenges of reinforcement learning. Springer
2020
Later among the works it cites.
Kim DK, Liu M, Omidshafiei S, Lopez-Cot S, Riemer M, Habibi G, Tesauro G, Mourad S, Campbell M, How JP (2020) Learning hierarchical teaching policies for cooperative agents. In: Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems, International Foundation for Autonomous Agents and Multi Agent Systems, Richland, SC, AAMAS ’20, p 620–628
2020
Later among the works it cites.
Lazaridou A, Baroni M (2020) Emergent multi-agent communication in the deep learning era. arXiv preprint arXiv:200602419
2020
Later among the works it cites.
Liu Z, Chen B, Zhou H, Koushik G, Hebert M, Zhao D (2020) Mapper: Multi-agent path planning with evolutionary reinforcement learning in mixed dynamic environments. In: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, pp 11748–11754
2020
Later among the works it cites.
Majumdar S, Khadka S, Miret S, Mcaleer S, Tumer K (2020) Evolutionary reinforcement learning for sample-efficient multiagent coordination. In: International Conference on Machine Learning, PMLR, pp 6651–6660
2020
Later among the works it cites.
Mao H, Zhang Z, Xiao Z, Gong Z, Ni Y (2020) Learning multi-agent communication with double attentional deep reinforcement learning. Autonomous Agents and Multi-Agent Systems 34(1):1–34
2020
Later among the works it cites.
McKee KR, Gemp I, McWilliams B, Duèñez Guzmán EA, Hughes E, Leibo JZ (2020) Social diversity and social preferences in mixed-motive reinforcement learning. In: Proceedings of the 19th International Conference on Autonomous Agents and Multiagent Systems, International Foundation for Autonomous Agents and Multi Agent Systems, Richland, SC, AAMAS ’20, p 869–877
2020
Later among the works it cites.
Palanisamy P (2020) Multi-agent connected autonomous driving using deep reinforcement learning. In: International Joint Conference on Neural Networks, IEEE, pp 1–7
2020
Later among the works it cites.
Plaat A (2020) Learning to Play: Reinforcement Learning and Games. Springer Nature
2020
Later among the works it cites.
Schrittwieser J, Antonoglou I, Hubert T, Simonyan K, Sifre L, Schmitt S, Guez A, Lockhart E, Hassabis D, Graepel T, et al. (2020) Mastering atari, go, chess and shogi by planning with a learned model. Nature 588(7839):604–609
2020
Later among the works it cites.
Sheikh HU, Bölöni L (2020) Multi-agent reinforcement learning for problems with combined individual and team reward. In: 2020 International Joint Conference on Neural Networks (IJCNN), IEEE, pp 1–8
2020
Later among the works it cites.
Zhou M, Liu Z, Sui P, Li Y, Chung YY (2020) Learning implicit credit assignment for cooperative multi-agent reinforcement learning. Advances in Neural Information Processing Systems 33:11853–11864
2020
Later among the works it cites.
Canese L, Cardarilli GC, Di Nunzio L, Fazzolari R, Giardino D, Re M, Spanò S (2021) Multi-agent reinforcement learning: A review of challenges and applications. Applied Sciences 11(11):4948
2021
Closest in time.
Castellini J, Devlin S, Oliehoek FA, Savani R (2021) Difference rewards policy gradients. In: Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems, International Foundation for Autonomous Agents and Multi Agent Systems, Richland, SC, AAMAS ’21, p 1475–1477
2021
Closest in time.
Cheng CA, Kolobov A, Swaminathan A (2021) Heuristic-guided reinforcement learning. Advances in Neural Information Processing Systems 34
2021
Closest in time.
Du W, Ding S (2021) A survey on multi-agent deep reinforcement learning: from the perspective of challenges and applications. Artificial Intelligence Review 54(5):3215–3238
2021
Closest in time.
Feriani A, Hossain E (2021) Single and multi-agent deep reinforcement learning for ai-enabled wireless networks: A tutorial. IEEE Communications Surveys & Tutorials 23(2):1226–1252
2021
Closest in time.
Gronauer S, Diepold K (2021) Multi-agent deep reinforcement learning: a survey. Artificial Intelligence Review pp 1–49
2021
Closest in time.
Gu S, Geng M, Lan L (2021) Attention-based fault-tolerant approach for multi-agent reinforcement learning systems. Entropy 23(9):1133
2021
Closest in time.
Hamrick JB, Friesen AL, Behbahani F, Guez A, Viola F, Witherspoon S, Anthony T, Buesing LH, Veličković P, Weber T (2021) On the role of planning in model-based deep reinforcement learning. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=IrM64DGB21
2021
Closest in time.
Levine S (2017) Berkeley CS 294-112, Lecture Notes: Model-Based Reinforcement Learning. URL: http://rail.eecs.berkeley.edu/deeprlcourse-fa17/f17docs/lecture_9_model_based_rl.pdf . Last visited on 2021/05/12
2021
Closest in time.
Ma Z, Luo Y, Ma H (2021) Distributed heuristic multi-agent path finding with communication. In: 2021 IEEE International Conference on Robotics and Automation (ICRA), IEEE, pp 8699–8705
2021
Closest in time.
Moreno P, Hughes E, McKee KR, Pires BA, Weber T (2021) Neural recursive belief states in multi-agent reinforcement learning. arXiv preprint arXiv:210202274
2021
Closest in time.
Su J, Adams S, Beling P (2021) Value-decomposition multi-agent actor-critics. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 35, pp 11352–11360
2021
Closest in time.
Taylor JET, Taylor GW (2021) Artificial cognition: How experimental psychology can help generate explainable artificial intelligence. Psychonomic Bulletin & Review 28(2):454–475
2021
Closest in time.
Terry JK, Grammel N, Hari A, Santos L, Black B (2021) Revisiting parameter sharing in multi-agent deep reinforcement learning. arXiv preprint arXiv:200513625
2021
Closest in time.
Tian R, Tomizuka M, Sun L (2021) Learning human rewards by inferring their latent intelligence levels in multi-agent games: A theory-of-mind approach with application to driving data. In: 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, pp 4560–4567
2021
Closest in time.
Wu Y, Mansimov E, Liao S, Radford A, Schulman J (2017b) openai baselines, acktr & a2c. http://https://openai.com/blog/baselines-acktr-a2c// , [Online; accessed 16-December-2021]
2021
Closest in time.
Yang Y, Wang J (2021) An overview of multi-agent reinforcement learning from game theoretical perspective. arXiv preprint arXiv:201100583
2021
Closest in time.
Zhang K, Yang Z, Başar T (2021) Multi-Agent Reinforcement Learning: A Selective Overview of Theories and Algorithms, Springer International Publishing, pp 321–384. DOI 10.1007/978-3-030-60990-0_12
2021
Closest in time.
Zou H, Ren T, Yan D, Su H, Zhu J (2021) Learning task-distribution reward shaping with meta-learning. In: Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, pp 2–9
2021
Closest in time.
Huang Y, Huang L, Zhu Q (2022) Reinforcement learning for feedback-enabled cyber resilience. Annual Reviews in Control
2022
Closest in time.
Peysakhovich A, Lerer A (2018) Prosocial learning agents solve generalized stag hunts better than selfish ones. In: International Foundation for Autonomous Agents and Multi Agent Systems, Richland, SC, AAMAS ’18, p 2043–2044
2044
Closest in time.
Sunehag P, Lever G, Gruslys A, Czarnecki WM, Zambaldi V, Jaderberg M, Lanctot M, Sonnerat N, Leibo JZ, Tuyls K, Graepel T (2018) Value-decomposition networks for cooperative multi-agent learning based on team reward. In: Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems, International Foundation for Autonomous Agents and Multi Agent Systems, Richland, SC, AAMAS ’18, p 2085–2087
2087
Closest in time.