Fetching the paper…
Reading the bibliography…
Deep Reinforcement Learning has made significant progress in multi-agent systems in recent years.
A survey on traffic signal control methods
Hua Wei, Guanjie Zheng, Vikash Gayah, and Zhenhui Li · 1904
Earlier work this paper cites.
A layered architecture for active perception: Image classification using deep reinforcement learning
Hossein K. Mousavi, Guangyi Liu, Weihang Yuan, Martin Takác, Héctor Muñoz-Avila, and Nader Motee · 1909
Earlier work this paper cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar · 1911
Earlier work this paper cites.
Colight: Learning network-level cooperation for traffic signal control
Hua Wei, Nan Xu, Huichu Zhang, Guanjie Zheng, Xinshi Zang, Chacha Chen, Weinan Zhang, Yanmin Zhu, Kai Xu, and Zhenhui Li · 1922
Earlier work this paper cites.
Methods of conjugate gradients for solving linear systems , volume 49
Magnus Rudolph Hestenes, Eduard Stiefel, et al · 1952
Earlier work this paper cites.
A markovian decision process
Richard Bellman · 1957
Earlier work this paper cites.
Studies in Linear and Non-linear Programming
Joseph Arrow Arrow, Leonid Hurwicz, and Hirofumi Uzawa · 1958
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
Distributed reinforcement learning
Gerhard Weiß · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Multiple path coordination for mobile robots: A geometric algorithm
Stephane Leroy, Jean-Paul Laumond, and Thierry Siméon · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart J. Russell · 1999
Earlier work this paper cites.
The o.d.e. method for convergence of stochastic approximation and reinforcement learning
Vivek S Borkar and Sean P Meyn · 2000
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
Martin Lauer and Martin Riedmiller · 2000
Earlier work this paper cites.
Multiagent systems: A survey from a machine learning perspective
Peter Stone and Manuela Veloso · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Hierarchical multi-agent reinforcement learning
Rajbala Makar, Sridhar Mahadevan, and Mohammad Ghavamzadeh · 2001
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
Daniel S Bernstein, Robert Givan, Neil Immerman, and Shlomo Zilberstein · 2002
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Michael Bowling and Manuela Veloso · 2002
Earlier work this paper cites.
Using a prm planner to compare centralized and decoupled planning for multi-robot systems
Gildardo Sanchez and J-C Latombe · 2002
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas · 2003
Earlier work this paper cites.
Unifying temporal and structural credit assignment problems
Adrian K Agogino and Kagan Tumer · 2004
Earlier work this paper cites.
Path planning in dynamic environments
Roman Smierzchalski and Zbigniew Michalewicz · 2005
Earlier work this paper cites.
Planning algorithms
Steven M LaValle · 2006
Earlier work this paper cites.
On a successful application of multi-agent reinforcement learning to operations research benchmarks
Thomas Gabel and Martin Riedmiller · 2007
Earlier work this paper cites.
A multiagent approach to q q -learning for daily stock trading
Jae Won Lee, Jonghun Park, O Jangmin, Jongwoo Lee, and Euyseok Hong · 2007
Earlier work this paper cites.
Hysteretic q-learning: an algorithm for decentralized reinforcement learning in cooperative multi-agent teams
Laëtitia Matignon, Guillaume Laurent, and Nadine Le Fort-Piat · 2007
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Lucian Bu, Robert Babu, Bart De Schutter, et al · 2008
Earlier work this paper cites.
Learning of communication codes in multi-agent reinforcement learning problem
Tatsuya Kasai, Hiroshi Tenmoto, and Akimoto Kamiya · 2008
Earlier work this paper cites.
Reinforcement learning in continuous action spaces through sequential monte carlo methods
Alessandro Lazaric, Marcello Restelli, and Andrea Bonarini · 2008
Earlier work this paper cites.
Convergent temporal-difference learning with arbitrary smooth function approximation
Shalabh Bhatnagar, Doina Precup, David Silver, Richard S Sutton, Hamid R Maei, and Csaba Szepesvári · 2009
Earlier work this paper cites.
A distributed actor-critic algorithm and applications to mobile sensor network coordination problems
Paris Pennesi and Ioannis Ch Paschalidis · 2009
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S. Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Earlier work this paper cites.
Efficient distributed reinforcement learning through agreement
Paulina Varshavskaya, Leslie Pack Kaelbling, and Daniela Rus · 2009
Earlier work this paper cites.
A multi-agent learning approach to online distributed resource allocation
Chongjie Zhang, Victor Lesser, and Prashant Shenoy · 2009
Earlier work this paper cites.
Reinforcement learning-based multi-agent system for network traffic signal control
Itamar Arel, Cong Liu, Tom Urbanik, and Airton G Kohls · 2010
Earlier work this paper cites.
Multi-agent reinforcement learning: An overview
Lucian Buşoniu, Robert Babuška, and Bart De Schutter · 2010
Earlier work this paper cites.
Double q-learning
Hado V Hasselt · 2010
Earlier work this paper cites.
Reinforcement learning with function approximation for traffic signal control
LA Prashanth and Shalabh Bhatnagar · 2010
Earlier work this paper cites.
Traffic light control in non-stationary environments based on multi agent q-learning
Monireh Abdoos, Nasser Mozayani, and Ana LC Bazzan · 2011
Earlier work this paper cites.
Theoretical considerations of potential-based reward shaping for multi-agent systems
Sam Devlin and Daniel Kudenko · 2011
Earlier work this paper cites.
Opensim: a musculoskeletal modeling and simulation framework for in silico investigations and exchange
Ajay Seth, Michael Sherman, Jeffrey A Reinbolt, and Scott L Delp · 2011
Earlier work this paper cites.
Reciprocal n-body collision avoidance
Jur Van Den Berg, Stephen J Guy, Ming Lin, and Dinesh Manocha · 2011
Earlier work this paper cites.
A novel multi-agent reinforcement learning approach for job scheduling in grid computing
Jun Wu, Xin Xu, Pengcheng Zhang, and Chunming Liu · 2011
Earlier work this paper cites.
Convergence of a multi-agent projected stochastic gradient algorithm for non-convex optimization
Pascal Bianchi and Jérémie Jakubowicz · 2012
Earlier work this paper cites.
Diffusion adaptation strategies for distributed optimization and learning over networks
Jianshu Chen and Ali H Sayed · 2012
Earlier work this paper cites.
Pareto-optimal coordination of multiple robots with safety guarantees
Rongxin Cui, Bo Gao, and Ji Guo · 2012
Earlier work this paper cites.
Monte carlo methods
Dirk P Kroese and Reuven Y Rubinstein · 2012
Earlier work this paper cites.
Transfer in reinforcement learning: a framework and a survey
Alessandro Lazaric · 2012
Earlier work this paper cites.
Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems
Laetitia Matignon, Guillaume J Laurent, and Nadine Le Fort-Piat · 2012
Earlier work this paper cites.
Termes: An autonomous robotic system for three-dimensional collective construction
Kirstin Petersen · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Asynchronous decentralized prioritized planning for coordination in multi-robot system
Michal Cáp, Peter Novák, Martin Seleckỳ, Jan Faigl, and Jiff Vokffnek · 2013
Earlier work this paper cites.
Distributed reinforcement learning in multi-agent networks
Soummya Kar, José MF Moura, and H Vincent Poor · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Suboptimal variants of the conflict-based search algorithm for the multi-agent pathfinding problem
Max Barer, Guni Sharon, Roni Stern, and Ariel Felner · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Potential-based difference rewards for multiagent reinforcement learning
Sam Devlin, Logan Yliniemi, Daniel Kudenko, and Kagan Tumer · 2014
Earlier work this paper cites.
Exploiting structure and agent-centric rewards to promote coordination in large multiagent systems
Chris HolmesParker, Matthew E. Taylor, Yusen Zhan, and Kagan Tumer · 2014
Earlier work this paper cites.
Distributed policy evaluation under multiple behavior strategies
Sergio Valcarcel Macua, Jianshu Chen, Santiago Zazo, and Ali H Sayed · 2014
Earlier work this paper cites.
Multi-agent reinforcement learning for traffic signal control
KJ Prabuchandran, Hemanth Kumar AN, and Shalabh Bhatnagar · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Empirically evaluating multiagent learning algorithms
Erik Zawadzki, Asher Lipson, and Kevin Leyton-Brown · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Matthew Hausknecht and Peter Stone · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Conflict-based search for optimal multi-agent pathfinding
Guni Sharon, Roni Stern, Ariel Felner, and Nathan R Sturtevant · 2015
Cited alongside, same era.
Mazebase: A sandbox for learning from games
Sainbayar Sukhbaatar, Arthur Szlam, Gabriel Synnaeve, Soumith Chintala, and Rob Fergus · 2015
Cited alongside, same era.
Subdimensional expansion for multirobot path planning
Primal-dual algorithm for distributed reinforcement learning: Distributed GTD
Donghwan Lee, Hyung-Jin Yoon, and Naira Hovakimyan · 2018
Later among the works it cites.
RLlib: Abstractions for distributed reinforcement learning
Eric Liang, Richard Liaw, Robert Nishihara, Philipp Moritz, Roy Fox, Ken Goldberg, Joseph Gonzalez, Michael Jordan, and Ion Stoica · 2018
Later among the works it cites.
Efficient large-scale fleet management via multi-agent deep reinforcement learning
Kaixiang Lin, Renyu Zhao, Zhe Xu, and Jiayu Zhou · 2018
Later among the works it cites.
The effects of memory replay in reinforcement learning
Ruishan Liu and James Zou · 2018
Later among the works it cites.
Diff-dac: Distributed actor-critic for average multitask deep reinforcement learning
Sergio Valcarcel Macua, Aleksi Tukiainen, Daniel García-Ocaña Hernández, David Baldazo, Enrique Munoz de Cote, and Santiago Zazo · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Glenn Wagner and Howie Choset · 2015
Cited alongside, same era.
A multi-agent framework for packet routing in wireless sensor networks
Dayong Ye, Minjie Zhang, and Yun Yang · 2015
Cited alongside, same era.
On convergence of emphatic temporal-difference learning
Huizhen Yu · 2015
Cited alongside, same era.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Cited alongside, same era.
Charles Beattie, Joel Z Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, et al · 2016
Cited alongside, same era.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Next: In-network nonconvex optimization
Paolo Di Lorenzo and Gesualdo Scutari · 2016
Cited alongside, same era.
David Mguni, Joel Jennings, Sergio Valcarcel Macua, Sofia Ceppi, and Enrique Munoz de Cote · 2018
Later among the works it cites.
Ray: A distributed framework for emerging { \{ AI } \} applications
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I Jordan, et al · 2018
Later among the works it cites.
Reinforcement learning for solving the vehicle routing problem
Mohammadreza Nazari, Afshin Oroojlooy, Lawrence Snyder, and Martin Takác · 2018
Later among the works it cites.
Machine theory of mind
Neil Rabinowitz, Frank Perbet, Francis Song, Chiyuan Zhang, S. M. Ali Eslami, and Matthew Botvinick · 2018
Later among the works it cites.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Later among the works it cites.
Multi-agent actor-critic with generative cooperative policy network
Heechang Ryu, Hayong Shin, and Jinkyoo Park · 2018
Later among the works it cites.
Learning when to communicate at scale in multiagent cooperative and competitive tasks
Amanpreet Singh, Tushar Jain, and Sainbayar Sukhbaatar · 2018
Later among the works it cites.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Hierarchical deep multiagent reinforcement learning
Hongyao Tang, Jianye Hao, Tangjie Lv, Yingfeng Chen, Zongzhang Zhang, Hangtian Jia, Chunxu Ren, Yan Zheng, Changjie Fan, and Li Wang · 2018
Later among the works it cites.
Multi-agent reinforcement learning via double averaging primal-dual optimization
Hoi-To Wai, Zhuoran Yang, Princeton Zhaoran Wang, and Mingyi Hong · 2018
Later among the works it cites.
Decentralised grid scheduling approach based on multi-agent reinforcement learning and gossip mechanism
Jun Wu and Xin Xu · 2018
Later among the works it cites.
Building generalizable agents with a realistic and rich 3d environment
Yi Wu, Yuxin Wu, Georgia Gkioxari, and Yuandong Tian · 2018
Later among the works it cites.
Ian Xiao · 2018
Later among the works it cites.
A finite sample analysis of the actor-critic algorithm
Zhuoran Yang, Kaiqing Zhang, Mingyi Hong, and Tamer Başar · 2018
Later among the works it cites.
Convergence of variance-reduced learning under random reshuffling
Bicheng Ying, Kun Yuan, and Ali H Sayed · 2018
Later among the works it cites.
Networked multi-agent reinforcement learning in continuous spaces
Kaiqing Zhang, Zhuoran Yang, and Tamer Basar · 2018
Later among the works it cites.
mazelab: A customizable framework to create maze and gridworld environments
Xingdong Zuo · 2018
Later among the works it cites.
Managing engineering systems with large state and action spaces through deep reinforcement learning
CP Andriotis and KG Papakonstantinou · 2019
Closest in time.
Multi-agent deep reinforcement learning for liquidation strategy analysis
Wenhang Bao and Xiao-yang Liu · 2019
Closest in time.
Autonomous air traffic controller: A deep multi-agent reinforcement learning approach
Marc Brittain and Peng Wei · 2019
Closest in time.
Multi-agent deep reinforcement learning for large-scale traffic signal control
Tianshu Chu, Jie Wang, Lara Codecà, and Zhaojian Li · 2019
Closest in time.
A survey on transfer learning for multiagent reinforcement learning systems
Felipe Leno Da Silva and Anna Helena Reali Costa · 2019
Closest in time.
TarMAC: Targeted multi-agent communication
Abhishek Das, Théophile Gervet, Joshua Romoff, Dhruv Batra, Devi Parikh, Mike Rabbat, and Joelle Pineau · 2019
Closest in time.
Actor-critic algorithms for constrained multi-agent reinforcement learning
Raghuram Bharadwaj Diddigi, Sai Koti Reddy Danda, Shalabh Bhatnagar, et al · 2019
Closest in time.
Reduced variance deep reinforcement learning with temporal logic specifications
Qitong Gao, Davood Hajinezhad, Yan Zhang, Yiannis Kantaros, and Michael M Zavlanos · 2019
Closest in time.
Decentralized network level adaptive signal control by multi-agent deep reinforcement learning
Yaobang Gong, Mohamed Abdel-Aty, Qing Cai, and Md Sharikur Rahman · 2019
Closest in time.
Actor-attention-critic for multi-agent reinforcement learning
Shariq Iqbal and Fei Sha · 2019
Closest in time.
Social influence as intrinsic motivation for multi-agent deep reinforcement learning
Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Caglar Gulcehre, Pedro Ortega, Dj Strouse, Joel Z. Leibo, and Nando De Freitas · 2019
Closest in time.
Multi-Agent Reinforcement Learning Environment
Shuo Jiang · 2019
Closest in time.
Learning to schedule communication in multi-agent reinforcement learning
Daewoo Kim, Sangwoo Moon, David Hostallero, Wan Ju Kang, Taeyoung Lee, Kyunghwan Son, and Yung Yi · 2019
Closest in time.
Neural trust region/proximal policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Closest in time.
Searching with consistent prioritization for multi-agent path finding
Hang Ma, Daniel Harabor, Peter J Stuckey, Jiaoyang Li, and Sven Koenig · 2019
Closest in time.
Modelling the dynamic joint policy of teammates with attention multi-agent ddpg
Hangyu Mao, Zhengchao Zhang, Zhen Xiao, and Zhibo Gong · 2019
Closest in time.
Multi-agent image classification via reinforcement learning
H. K. Mousavi, M. Nazari, M. Takáč, and N. Motee · 2019
Closest in time.
The StarCraft Multi-Agent Challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philiph H. S. Torr, Jakob Foerster, and Shimon Whiteson · 2019
Closest in time.
Multi-agent common knowledge reinforcement learning
Christian Schroeder de Witt, Jakob Foerster, Gregory Farquhar, Philip Torr, Wendelin Boehmer, and Shimon Whiteson · 2019
Closest in time.
M 3 RL: Mind-aware multi-agent management reinforcement learning
Tianmin Shu and Yuandong Tian · 2019
Closest in time.
A reinforcement learning-based multi-agent framework applied for solving routing and scheduling problems
Maria Amélia Lopes Silva, Sérgio Ricardo de Souza, Marcone Jamilson Freitas Souza, and Ana Lúcia C Bazzan · 2019
Closest in time.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, , and Yung Yi · 2019
Closest in time.
Neural mmo: A massively multiagent game environment for training and evaluating intelligent agents
Joseph Suarez, Yilun Du, Phillip Isola, and Igor Mordatch · 2019
Closest in time.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Closest in time.
R-maddpg for partially observable environments and limited communication
Rose E Wang, Michael Everett, and Jonathan P How · 2019
Closest in time.
Distributed off-policy actor-critic reinforcement learning with policy consensus
Y. Zhang and M. M. Zavlanos · 2019
Closest in time.
Learning phase competition for traffic signal control
Guanjie Zheng, Yuanhao Xiong, Xinshi Zang, Jie Feng, Hua Wei, Huichu Zhang, Yong Li, Kai Xu, and Zhenhui Li · 2019
Closest in time.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2020
Closest in time.
Cooperative zone-based rebalancing of idle overhead hoist transportations using multi-agent reinforcement learning with graph representation learning
Kyuree Ahn and Jinkyoo Park · 2020
Closest in time.
The hanabi challenge: A new frontier for ai research
Nolan Bard, Jakob N Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H Francis Song, Emilio Parisotto, Vincent Dumoulin, Subhodeep Moitra, Edward Hughes, et al · 2020
Closest in time.
Multiagent fully decentralized value function learning with linear convergence rates
L. Cassano, K. Yuan, and A. H. Sayed · 2020
Closest in time.
Toward a thousand lights: Decentralized deep reinforcement learning for large-scale traffic signal control
Hua Wei Chacha Chen, Nan Xu, Guanjie Zheng, Ming Yang, Yuanhao Xiong, Kai Xu, and Zhenhui Li · 2020
Closest in time.
Expected policy gradients for reinforcement learning
Kamil Ciosek and Shimon Whiteson · 2020
Closest in time.
Cooperative multi-agent system for production control using reinforcement learning
Marc-André Dittrich and Silas Fohlmeister · 2020
Closest in time.
Revisiting fundamentals of experience replay
William Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio, Hugo Larochelle, Mark Rowland, and Will Dabney · 2020
Closest in time.
Communication learning via backpropagation in discrete channels with unknown noise
Benjamin Freed, Guillaume Sartoretti, Jiaheng Hu, and Howie Choset · 2020
Closest in time.
Graph convolutional reinforcement learning
Jiechuan Jiang, Chen Dun, Tiejun Huang, and Zongqing Lu · 2020
Closest in time.
Model-based reinforcement learning: A survey
Thomas M Moerland, Joost Broekens, and Catholijn M Jonker · 2020
Closest in time.
Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications
Thanh Thi Nguyen, Ngoc Duy Nguyen, and Saeid Nahavandi · 2020
Closest in time.
A deep q-network for the beer game: Deep reinforcement learning for inventory optimization
Afshin Oroojlooyjadid, MohammadReza Nazari, Lawrence V. Snyder, and Martin Takáč · 2020
Closest in time.
Reinforcement learning with dynamic boltzmann softmax updates
Ling Pan, Qingpeng Cai, Qi Meng, Wei Chen, and Longbo Huang · 2020
Closest in time.
Arena: A general evaluation platform and building toolkit for multi-agent intelligence
Yuhang Song, Andrzej Wojcicki, Thomas Lukasiewicz, Jianyi Wang, Abi Aryan, Zhenghua Xu, Mai Xu, Zihan Ding, and Lianlong Wu · 2020
Closest in time.
Cm3: Cooperative multi-goal multi-stage multi-agent reinforcement learning
Jiachen Yang, Alireza Nakhaei, David Isele, Kikuo Fujimura, and Hongyuan Zha · 2020
Closest in time.
An overview of multi-agent reinforcement learning from game theoretical perspective
Yaodong Yang and Jun Wang · 2020
Closest in time.
A survey of inverse reinforcement learning: Challenges, methods and progress
Saurabh Arora and Prashant Doshi · 2021
Closest in time.
Modelling stock markets by multi-agent reinforcement learning
Johann Lussange, Ivan Lazarevich, Sacha Bourgeois-Gironde, Stefano Palminteri, and Boris Gutkin · 2021
Closest in time.
A multi-agent off-policy actor-critic algorithm for distributed reinforcement learning
Wesley Suttle, Zhuoran Yang, Kaiqing Zhang, Zhaoran Wang, Tamer Başar, and Ji Liu · 2021
Closest in time.