Fetching the paper…
Reading the bibliography…
Multi-agent systems (MAS) are widely prevalent and crucially important in numerous real-world applications, where multiple agents must make decisions to achieve their objectives in a shared environment.
Gpu-accelerated atari emulation for reinforcement learning
S. Dalton, I. Frosio, and M. Garland · 1907
Earlier work this paper cites.
Non-cooperative games
J. F. Nash · 1951
Earlier work this paper cites.
An iterative method of solving a game
J. J. Robinson · 1951
Earlier work this paper cites.
A value for n-person games
L. S. Shapley · 1952
Earlier work this paper cites.
Stochastic games*
L. S. Shapley · 1953
Earlier work this paper cites.
An analog of the minimax theorem for vector payoffs
D. Blackwell · 1956
Earlier work this paper cites.
Dynamic programming princeton university press princeton
R. Bellman · 1957
Earlier work this paper cites.
La reconstruction du nid et les coordinations interindividuelles chez bellicositermes natalensis et cubitermes sp. la théorie de la stigmergie: Essai d’interprétation du comportement des termites constructeurs
P.-P. Grassé · 1959
Earlier work this paper cites.
Games with incomplete information played by “bayesian” players, i–iii part i. the basic model
J. C. Harsanyi · 1967
Earlier work this paper cites.
Subjectivity and correlation in randomized strategies
R. J. Aumann · 1974
Earlier work this paper cites.
The extragradient method for finding saddle points and other problems
G. M. Korpelevich · 1976
Earlier work this paper cites.
Evolution and the theory of games
J. Maynard Smith · 1976
Earlier work this paper cites.
Does the chimpanzee have a theory of mind?
D. Premack and G. Woodruff · 1978
Earlier work this paper cites.
Utility theory for decision making
P. C. Fishburn, P. C. Fishburn, et al · 1979
Earlier work this paper cites.
Sequential equilibria
D. M. Kreps and R. Wilson · 1982
Earlier work this paper cites.
An introduction to game theory
R. B. Myerson · 1985
Earlier work this paper cites.
On the concept of dynamic multi-level simulation
S. Ghosh · 1986
Earlier work this paper cites.
Reexamination of the perfectness concept for equilibrium points in extensive games, 1988
R. S. Bielefeld · 1988
Earlier work this paper cites.
Toward a theory of reinforcement-learning connectionist systems
R. J. Williams · 1988
Earlier work this paper cites.
Nash and correlated equilibria: Some complexity considerations
I. Gilboa and E. Zemel · 1989
Earlier work this paper cites.
Learning from delayed rewards
C. Watkins · 1989
Earlier work this paper cites.
Finite-dimensional variational inequality and nonlinear complementarity problems: a survey of theory, algorithms and applications
P. T. Harker and J.-S. Pang · 1990
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
R. S. Sutton · 1990
Earlier work this paper cites.
Bayesian learning in normal form games
J. Jordan · 1991
Earlier work this paper cites.
Advances in prospect theory: Cumulative representation of uncertainty
A. Tversky and D. Kahneman · 1992
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
P. Dayan · 1993
Earlier work this paper cites.
Learning mixed equilibria
D. Fudenberg and D. M. Kreps · 1993
Earlier work this paper cites.
Rational learning leads to nash equilibrium
E. Kalai and E. Lehrer · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Learning to behave socially
M. J. Matarić · 1994
Earlier work this paper cites.
Incorporating opponent models into adversary search
D. Carmel and S. Markovitch · 1996
Earlier work this paper cites.
Progress in behavioral game theory
C. F. Camerer · 1997
Earlier work this paper cites.
Reinforcement learning in the multi-robot domain
M. J. Matarić · 1997
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents, 1997
M. Tan · 1997
Earlier work this paper cites.
Dynamic noncooperative game theory
T. Başar and G. J. Olsder · 1998
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
C. Claus and C. Boutilier · 1998
Earlier work this paper cites.
An environment model for nonstationary reinforcement learning
S. Choi, D.-Y. Yeung, and N. Zhang · 1999
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Y. Freund and R. E. Schapire · 1999
Earlier work this paper cites.
Stigmergy, self-organization, and sorting in collective robotics
O. Holland and C. Melhuish · 1999
Earlier work this paper cites.
Worst-case equilibria
E. Koutsoupias and C. Papadimitriou · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
A brief history of stigmergy
G. Theraulaz and E. Bonabeau · 1999
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
S. Hart and A. Mas-Colell · 2000
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
M. Lauer and M. A. Riedmiller · 2000
Earlier work this paper cites.
Nash convergence of gradient dynamics in general-sum games
S. Singh, M. Kearns, and Y. Mansour · 2000
Earlier work this paper cites.
Multiagent systems: A survey from a machine learning perspective
P. Stone and M. Veloso · 2000
Earlier work this paper cites.
Sociobiology
E. O. Wilson · 2000
Earlier work this paper cites.
An analysis of stochastic game theory for multiagent reinforcement learning
M. Bowling and M. Veloso · 2001
Earlier work this paper cites.
Multiagent planning with factored mdps
C. Guestrin, D. Koller, and R. Parr · 2001
Earlier work this paper cites.
A natural policy gradient
S. M. Kakade · 2001
Earlier work this paper cites.
Optimal payoff functions for members of collectives
D. H. Wolpert and K. Tumer · 2001
Earlier work this paper cites.
Multiagent learning using a variable learning rate
M. Bowling and M. Veloso · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. M. Kakade and J. Langford · 2002
Earlier work this paper cites.
Learning to coordinate efficiently: A model-based approach
R. I. Brafman and M. Tennenholtz · 2003
Earlier work this paper cites.
Coordination in multiagent reinforcement learning: A bayesian approach
G. Chalkiadakis and C. Boutilier · 2003
Earlier work this paper cites.
Awesome: A general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents, 2003
V. Conitzer and T. Sandholm · 2003
Earlier work this paper cites.
Correlated-q learning
A. Greenwald and K. Hall · 2003
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
J. Hu and M. P. Wellman · 2003
Earlier work this paper cites.
Improving on the reinforcement learning of coordination in cooperative multi-agent systems
S. Kapetanakis and D. Kudenko · 2003
Earlier work this paper cites.
Multi-agent influence diagrams for representing and solving games
D. Koller and B. Milch · 2003
Earlier work this paper cites.
Do anomalies disappear in repeated markets?
G. Loomes, C. Starmer, and R. Sugden · 2003
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
H. B. McMahan, G. J. Gordon, and A. Blum · 2003
Earlier work this paper cites.
Multi-agent reinforcement learning:a critical survey
Y. Shoham, R. Powers, and T. Grenager · 2003
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
Unifying temporal and structural credit assignment problems
A. K. Agogino and K. Tumer · 2004
Earlier work this paper cites.
Convergence and no-regret in multiagent learning
M. Bowling · 2004
Earlier work this paper cites.
Anytime algorithms for multiagent decision making using coordination graphs
N. Vlassis, R. Elhorst, and J. Kok · 2004
Earlier work this paper cites.
Best-response multiagent learning in non-stationary environments
M. Weinberg and J. Rosenschein · 2004
Earlier work this paper cites.
Ant colony optimization theory: A survey
M. Dorigo and C. Blum · 2005
Earlier work this paper cites.
Utile coordination: Learning interdependencies among cooperative agents
J. R. Kok, E. J. Hoen, B. Bakker, and N. Vlassis · 2005
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2005
Earlier work this paper cites.
Networked distributed pomdps: A synthesis of distributed constraint optimization and pomdps
R. Nair, P. Varakantham, M. Tambe, and M. Yokoo · 2005
Earlier work this paper cites.
Cooperative multi-agent learning: The state of the art
L. Panait and S. Luke · 2005
Earlier work this paper cites.
Beyond the centralized mindset–explorations in massively-parallel microworlds
M. Resnick · 2005
Earlier work this paper cites.
The economics of contracts: a primer
B. Salanié · 2005
Earlier work this paper cites.
Model compression
C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil · 2006
Earlier work this paper cites.
Designing economic mechanisms
L. Hurwicz and S. Reiter · 2006
Earlier work this paper cites.
If multi-agent learning is the answer, what is the question?
Y. Shoham, R. Powers, and T. Grenager · 2006
Earlier work this paper cites.
An evolutionary dynamical analysis of multi-agent learning in iterated games
K. Tuyls, P. J. T. Hoen, and B. Vanschoenwinkel · 2006
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
L. Busoniu, R. Babuska, and B. De Schutter · 2007
Earlier work this paper cites.
Collision avoidance in multi-robot systems
C. Cai, C. Yang, Q. Zhu, and Y. Liang · 2007
Earlier work this paper cites.
Settling the complexity of computing two-player nash equilibria, 2007
X. Chen, X. Deng, and S.-H. Teng · 2007
Earlier work this paper cites.
Predicting and preventing coordination problems in cooperative q-learning systems
N. Fulda and D. Ventura · 2007
Earlier work this paper cites.
Hysteretic q-learning : an algorithm for decentralized reinforcement learning in cooperative multi-agent teams
L. Matignon, G. J. Laurent, and N. Le Fort-Piat · 2007
Earlier work this paper cites.
Algorithmic game theory
N. Nisan, T. Roughgarden, E. Tardos, and V. V. Vazirani · 2007
Earlier work this paper cites.
Distributed agent-based air traffic flow management
K. Tumer and A. Agogino · 2007
Earlier work this paper cites.
Regret minimization in games with incomplete information
M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione · 2007
Earlier work this paper cites.
Analyzing and visualizing multiagent rewards in dynamic and stochastic domains
A. K. Agogino and K. Tumer · 2008
Earlier work this paper cites.
The price of stability for network design with fair cost allocation
E. Anshelevich, A. Dasgupta, J. Kleinberg, É. Tardos, T. Wexler, and T. Roughgarden · 2008
Earlier work this paper cites.
New complexity results about nash equilibria
V. Conitzer and T. Sandholm · 2008
Earlier work this paper cites.
Convention: A philosophical study
D. Lewis · 2008
Earlier work this paper cites.
Stigmergic epistemology, stigmergic cognition
L. Marsh and C. Onof · 2008
Earlier work this paper cites.
Multiagent systems: Algorithmic, game-theoretic, and logical foundations
Y. Shoham and K. Leyton-Brown · 2008
Earlier work this paper cites.
The complexity of computing a nash equilibrium
C. Daskalakis, P. W. Goldberg, and C. H. Papadimitriou · 2009
Earlier work this paper cites.
The theory of incentives: the principal-agent model
J.-J. Laffont and D. Martimort · 2009
Earlier work this paper cites.
Learning of coordination: Exploiting sparse interactions in multiagent systems
F. S. Melo and M. Veloso · 2009
Earlier work this paper cites.
An introduction to game theory
M. J. Osborne · 2009
Earlier work this paper cites.
Multi-agent reinforcement learning: An overview
L. Busoniu, R. Babuvska, and B. D. Schutter · 2010
Earlier work this paper cites.
Learning multi-agent state space representations
Y.-M. De Hauwere, P. Vrancx, and A. Nowé · 2010
Earlier work this paper cites.
On a connection between importance sampling and the likelihood ratio policy gradient
T. Jie and P. Abbeel · 2010
Earlier work this paper cites.
Using aseme methodology for model-driven agent systems development
N. Spanoudakis and P. Moraitis · 2010
Earlier work this paper cites.
Ad hoc autonomous agent teams: Collaboration without pre-coordination
P. Stone, G. A. Kaminka, S. Kraus, and J. S. Rosenschein · 2010
Earlier work this paper cites.
Rode: Learning roles to decompose multi-agent tasks
T. Wang, T. Gupta, A. Mahajan, B. Peng, S. Whiteson, and C. Zhang · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Action selection via learning behavior patterns in multi-robot domains
C. Erdogan and M. Veloso · 2011
Earlier work this paper cites.
Action-graph games
A. X. Jiang, K. Leyton-Brown, and N. A. Bhat · 2011
Earlier work this paper cites.
Decentralized mdps with sparse interactions
F. S. Melo and M. Veloso · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning, 2011
S. Ross, G. J. Gordon, and J. A. Bagnell · 2011
Cited alongside, same era.
The multiplicative weights update method: a meta-algorithm and applications
S. Arora, E. Hazan, and S. Kale · 2012
Cited alongside, same era.
T. Degris, M. White, and R. S. Sutton · 2012
Cited alongside, same era.
Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems
L. Matignon, G. J. Laurent, and N. Le Fort-Piat · 2012
Cited alongside, same era.
Online learning and online convex optimization
S. Shalev-Shwartz et al · 2012
Cited alongside, same era.
Multiagent learning: Basics, challenges, and prospects
Learning to schedule communication in multi-agent reinforcement learning, 2019
D. Kim, S. Moon, D. Hostallero, W. J. Kang, T. Lee, K. Son, and Y. Yi · 2019
Later among the works it cites.
ma-gym: Collection of multi-agent environments based on openai gym
A. Koul · 2019
Later among the works it cites.
On the pitfalls of measuring emergent communication, 2019
R. Lowe, J. Foerster, Y.-L. Boureau, J. Pineau, and Y. Dauphin · 2019
Later among the works it cites.
Applications of deep reinforcement learning in communications and networking: A survey
N. C. Luong, D. T. Hoang, S. Gong, D. Niyato, P. Wang, Y.-C. Liang, and D. I. Kim · 2019
Later among the works it cites.
A unified analysis of extra-gradient and optimistic gradient methods for saddle point problems: Proximal point approach, 2019
A. Mokhtari, A. Ozdaglar, and S. Pattathil · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Tuyls and G. Weiss · 2012
Cited alongside, same era.
The complexity of decentralized control of markov decision processes
D. S. Bernstein, S. Zilberstein, and N. Immerman · 2013
Cited alongside, same era.
An introduction to statistical learning , volume 112
G. James, D. Witten, T. Hastie, R. Tibshirani, et al · 2013
Cited alongside, same era.
Prospect theory: An analysis of decision under risk
D. Kahneman and A. Tversky · 2013
Cited alongside, same era.
Monte Carlo sampling and regret minimization for equilibrium computation and decision-making in large extensive form games
M. Lanctot · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Cited alongside, same era.
Optimization, learning, and games with predictable sequences, 2013
A. Rakhlin and K. Sridharan · 2013
Cited alongside, same era.
G. Palmer, R. Savani, and K. Tuyls · 2019
Later among the works it cites.
Dealing with non-stationarity in multi-agent deep reinforcement learning, 2019
G. Papoudakis, F. Christianos, A. Rahman, and S. V. Albrecht · 2019
Later among the works it cites.
A tour of reinforcement learning: The view from continuous control
B. Recht · 2019
Later among the works it cites.
The StarCraft Multi-Agent Challenge
M. Samvelyan, T. Rashid, C. S. de Witt, G. Farquhar, N. Nardelli, T. G. J. Rudner, C.-M. Hung, P. H. S. Torr, J. Foerster, and S. Whiteson · 2019
Later among the works it cites.
A survey on transfer learning for multiagent reinforcement learning systems
F. Silva and A. Costa · 2019
Later among the works it cites.
Fully parameterized quantile function for distributional reinforcement learning
D. Yang, L. Zhao, Z. Lin, T. Qin, J. Bian, and T.-Y. Liu · 2019
Later among the works it cites.
Learning to communicate in multi-agent reinforcement learning : A review, 2019
M. S. Zaïem and E. Bennequin · 2019
Later among the works it cites.
Solar: Deep structured representations for model-based reinforcement learning
M. Zhang, S. Vikram, L. Smith, P. Abbeel, M. Johnson, and S. Levine · 2019
Later among the works it cites.
Deep coordination graphs, 2020
W. Böhmer, V. Kurin, and S. Whiteson · 2020
Later among the works it cites.
On the utility of learning about humans for human-ai coordination, 2020
M. Carroll, R. Shah, M. K. Ho, T. L. Griffiths, S. A. Seshia, P. Abbeel, and A. Dragan · 2020
Later among the works it cites.
Aateam: Achieving the ad hoc teamwork by employing the attention mechanism
S. Chen, E. Andrejczuk, Z. Cao, and J. Zhang · 2020
Later among the works it cites.
Shared experience actor-critic for multi-agent reinforcement learning
F. Christianos, L. Schäfer, and S. V. Albrecht · 2020
Later among the works it cites.
Multi-agent reinforcement learning for networked system control, 2020
T. Chu, S. Chinchali, and S. Katti · 2020
Later among the works it cites.
Phasic policy gradient, 2020
K. Cobbe, J. Hilton, O. Klimov, and J. Schulman · 2020
Later among the works it cites.
Uncertainty-aware action advising for deep reinforcement learning agents
F. L. Da Silva, P. Hernandez-Leal, B. Kartal, and M. E. Taylor · 2020
Later among the works it cites.
Communication learning via backpropagation in discrete channels with unknown noise
B. Freed, G. Sartoretti, J. Hu, and H. Choset · 2020
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination, 2020
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2020
Later among the works it cites.
G. Hu, Y. Zhu, D. Zhao, M. Zhao, and J. Hao · 2020
Later among the works it cites.
Extragradient with player sampling for faster nash equilibrium finding
S. Jelassi, C. Domingo-Enrich, D. Scieur, A. Mensch, and J. Bruna · 2020
Later among the works it cites.
Graph convolutional reinforcement learning, 2020
J. Jiang, C. Dun, T. Huang, and Z. Lu · 2020
Later among the works it cites.
Communication in multi-agent reinforcement learning: Intention sharing
W. Kim, J. Park, and Y. Sung · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning, 2020
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Google research football: A novel reinforcement learning environment, 2020
K. Kurach, A. Raichuk, P. Stańczyk, M. Zając, O. Bachem, L. Espeholt, C. Riquelme, D. Vincent, M. Michalski, O. Bousquet, and S. Gelly · 2020
Later among the works it cites.
K.-H. Lai, D. Zha, Y. Li, and X. Hu · 2020
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments, 2020
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch · 2020
Later among the works it cites.
Likelihood quantile networks for coordinating multi-agent reinforcement learning, 2020
X. Lyu and C. Amato · 2020
Later among the works it cites.
Maven: Multi-agent variational exploration, 2020
A. Mahajan, T. Rashid, M. Samvelyan, and S. Whiteson · 2020
Later among the works it cites.
Independent Learning Approaches: Overcoming Multi-Agent Learning Pathologies In Team-Games
G. Palmer · 2020
Later among the works it cites.
Factored value functions for cooperative multi-agent reinforcement learning
S. Whiteson · 2020
Later among the works it cites.
Smarts: Scalable multi-agent reinforcement learning training school for autonomous driving, 2020
M. Zhou, J. Luo, J. Villella, Y. Yang, D. Rusu, J. Miao, W. Zhang, M. Alban, I. Fadakar, Z. Chen, A. C. Huang, Y. Wen, K. Hassanzadeh, D. Graves, D. Chen, Z. Zhu, N. Nguyen, M. Elsayed, K. Shao, S. Ahilan, B. Zhang, J. Wu, Z. Fu, K. Rezaee, P. Yadmellat, M. Rohani, N. P. Nieves, Y. Ni, S. Banijamali, A. C. Rivers, Z. Tian, D. Palenicek, H. bou Ammar, H. Zhang, W. Liu, J. Hao, and J. Wang · 2020
Later among the works it cites.
Multi-agent safe planning with gaussian processes
Z. Zhu, E. Biyik, and D. Sadigh · 2020
Later among the works it cites.
Double oracle algorithm for computing equilibria in continuous games
L. Adam, R. Horčík, T. Kasl, and T. Kroupa · 2021
Later among the works it cites.
Multi-agent inverse reinforcement learning: Suboptimal demonstrations and alternative solution concepts, 2021
S. Bergerson · 2021
Later among the works it cites.
Learning successor states and goal-dependent values: A mathematical viewpoint, 2021
L. Blier, C. Tallec, and Y. Ollivier · 2021
Later among the works it cites.
Brax – a differentiable physics engine for large scale rigid body simulation, 2021
C. D. Freeman, E. Frey, A. Raichuk, S. Girgin, I. Mordatch, and O. Bachem · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning, 2021
S. Fujimoto and S. S. Gu · 2021
Later among the works it cites.
Off-belief learning, 2021
H. Hu, A. Lerer, B. Cui, D. Wu, L. Pineda, N. Brown, and J. Foerster · 2021
Later among the works it cites.
Mix and mask actor-critic methods
D. Huh · 2021
Later among the works it cites.
Coordinated exploration via intrinsic rewards for multi-agent reinforcement learning, 2021
S. Iqbal and F. Sha · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning, 2021
I. Kostrikov, A. Nair, and S. Levine · 2021
Later among the works it cites.
Warpdrive: Extremely fast end-to-end deep multi-agent reinforcement learning on a gpu, 2021
T. Lan, S. Srinivasa, H. Wang, and S. Zheng · 2021
Later among the works it cites.
Deep implicit coordination graphs for multi-agent reinforcement learning, 2021
S. Li, J. K. Gupta, P. Morales, R. Allen, and M. J. Kochenderfer · 2021
Later among the works it cites.
Learning to ground multi-agent communication with autoencoders
T. Lin, J. Huh, C. Stauffer, S. N. Lim, and P. Isola · 2021
Later among the works it cites.
Cooperative exploration for multi-agent deep reinforcement learning, 2021
I.-J. Liu, U. Jain, R. A. Yeh, and A. G. Schwing · 2021
Later among the works it cites.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees, 2021
Y. Luo, H. Xu, Y. Li, Y. Tian, T. Darrell, and T. Ma · 2021
Later among the works it cites.
Conservative offline distributional reinforcement learning
Y. Ma, D. Jayaraman, and O. Bastani · 2021
Later among the works it cites.
Isaac gym: High performance gpu-based physics simulation for robot learning, 2021
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, and G. State · 2021
Later among the works it cites.
Xdo: A double oracle algorithm for extensive-form games
S. McAleer, J. B. Lanier, K. A. Wang, P. Baldi, and R. Fox · 2021
Later among the works it cites.
Emergent social learning via multi-agent reinforcement learning
K. K. Ndousse, D. Eck, S. Levine, and N. Jaques · 2021
Later among the works it cites.
Scalable reinforcement learning for multi-agent networked systems, 2021
G. Qu, A. Wierman, and N. Li · 2021
Later among the works it cites.
Ad hoc teamwork in the presence of non-stationary teammates
P. Santos, J. Ribeiro, A. Sardinha, and F. Melo · 2021
Later among the works it cites.
Dfac framework: Factorizing the value function via quantile mixture for multi-agent distributional q-learning, 2021
W.-F. Sun, C.-K. Lee, and C.-Y. Lee · 2021
Later among the works it cites.
Pettingzoo: Gym for multi-agent reinforcement learning
J. Terry, B. Black, N. Grammel, M. Jayakumar, A. Hari, R. Sullivan, L. S. Santos, C. Dieffendahl, C. Horsch, R. Perez-Vicente, et al · 2021
Later among the works it cites.
Learning one representation to optimize all rewards, 2021
A. Touati and Y. Ollivier · 2021
Later among the works it cites.
A new formalism, method and open issues for zero-shot coordination
J. Treutlein, M. Dennis, C. Oesterheld, and J. Foerster · 2021
Later among the works it cites.
Linear last-iterate convergence in constrained saddle-point optimization, 2021
C.-Y. Wei, C.-W. Lee, M. Zhang, and H. Luo · 2021
Later among the works it cites.
Coordination between individual agents in multi-agent reinforcement learning
Y. Zhang, Q. Yang, D. An, and C. Zhang · 2021
Later among the works it cites.
Episodic multi-agent reinforcement learning with curiosity-driven exploration, 2021
L. Zheng, J. Chen, J. Wang, J. He, Y. Hu, Y. Chen, C. Fan, Y. Gao, and C. Zhang · 2021
Later among the works it cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos, 2022
B. Baker, I. Akkaya, P. Zhokhov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Later among the works it cites.
Vmas: A vectorized multi-agent simulator for collective robot learning
M. Bettini, R. Kortvelesy, J. Blumenkamp, and A. Prorok · 2022
Later among the works it cites.
On the opportunities and risks of foundation models, 2022
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolfsson, S. Buch, D. Card, R. Castellon, N. Chatterji, A. Chen, K. Creel, J. Q. Davis, D. Demszky, C. Donahue, M. Doumbouya, E. Durmus, S. Ermon, J. Etchemendy, K. Ethayarajh, L. Fei-Fei, C. Finn, T. Gale, L. Gillespie, K. Goel, N. Goodman, S. Grossman, N. Guha, T. Hashimoto, P. Henderson, J. Hewitt, D. E. Ho, J. Hong, K. Hsu, J. Huang, T. Icard, S. Jain, D. Jurafsky, P. Kalluri, S. Karamcheti, G. Keeling, F. Khani, O. Khattab, P. W. Koh, M. Krass, R. Krishna, R. Kuditipudi, A. Kumar, F. Ladhak, M. Lee, T. Lee, J. Leskovec, I. Levent, X. L. Li, X. Li, T. Ma, A. Malik, C. D. Manning, S. Mirchandani, E. Mitchell, Z. Munyikwa, S. Nair, A. Narayan, D. Narayanan, B. Newman, A. Nie, J. C. Niebles, H. Nilforoshan, J. Nyarko, G. Ogut, L. Orr, I. Papadimitriou, J. S. Park, C. Piech, E. Portelance, C. Potts, A. Raghunathan, R. Reich, H. Ren, F. Rong, Y. Roohani, C. Ruiz, J. Ryan, C. Ré, D. Sadigh, S. Sagawa, K. Santhanam, A. Shih, K. Srinivasan, A. Tamkin, R. Taori, A. W. Thomas, F. Tramèr, R. E. Wang, W. Wang, B. Wu, J. Wu, Y. Wu, S. M. Xie, M. Yasunaga, J. You, M. Zaharia, M. Zhang, T. Zhang, X. Zhang, Y. Zhang, L. Zheng, K. Zhou, and P. Liang · 2022
Later among the works it cites.
How attentive are graph attention networks?, 2022
S. Brody, U. Alon, and E. Yahav · 2022
Later among the works it cites.
Towards human-level bimanual dexterous manipulation with reinforcement learning
Y. Chen, Y. Yang, T. Wu, S. Wang, X. Feng, J. Jiang, Z. Lu, S. M. McAleer, H. Dong, and S.-C. Zhu · 2022
Later among the works it cites.
Equilibrium computation and machine learning, 2022
C. Daskalakis · 2022
Later among the works it cites.
The complexity of markov equilibrium in stochastic games, 2022
C. Daskalakis, N. Golowich, and K. Zhang · 2022
Later among the works it cites.
Smacv2: An improved benchmark for cooperative multi-agent reinforcement learning, 2022
B. Ellis, S. Moalla, M. Samvelyan, M. Sun, A. Mahajan, J. N. Foerster, and S. Whiteson · 2022
Later among the works it cites.
Revisiting some common practices in cooperative multi-agent reinforcement learning, 2022
W. Fu, C. Yu, Z. Xu, J. Yang, and Y. Wu · 2022
Later among the works it cites.
Multi-agent deep reinforcement learning: A survey
S. Gronauer and K. Diepold · 2022
Later among the works it cites.
Mastering atari with discrete world models, 2022
D. Hafner, T. Lillicrap, M. Norouzi, and J. Ba · 2022
Later among the works it cites.
Human-level atari 200x faster, 2022
S. Kapturowski, V. Campos, R. Jiang, N. Rakićević, H. van Hasselt, C. Blundell, and A. P. Badia · 2022
Later among the works it cites.
Influencing long-term behavior in multiagent reinforcement learning, 2022
D.-K. Kim, M. Riemer, M. Liu, J. N. Foerster, M. Everett, C. Sun, G. Tesauro, and J. P. How · 2022
Later among the works it cites.
Difference advantage estimation for multi-agent policy gradients
Y. Li, G. Xie, and Z. Lu · 2022
Later among the works it cites.
A survey of ad hoc teamwork research, 2022
R. Mirsky, I. Carlucho, A. Rahman, E. Fosong, W. Macke, M. Sridharan, P. Stone, and S. V. Albrecht · 2022
Later among the works it cites.
A survey of opponent modeling in adversarial domains
S. Nashed and S. Zilberstein · 2022
Later among the works it cites.
Plan better amid conservatism: Offline multi-agent reinforcement learning with actor rectification
L. Pan, L. Huang, T. Ma, and H. Xu · 2022
Later among the works it cites.
Risk-sensitive reinforcement learning via policy gradient search
L. Prashanth, M. C. Fu, et al · 2022
Later among the works it cites.
Self-organized group for cooperative multi-agent reinforcement learning
J. Shao, Z. Lou, H. Zhang, Y. Jiang, S. He, and X. Ji · 2022
Later among the works it cites.
Offline multi-agent reinforcement learning with knowledge distillation
W.-C. Tseng, T.-H. J. Wang, Y.-C. Lin, and P. Isola · 2022
Later among the works it cites.
Individual reward assisted multi-agent reinforcement learning
L. Wang, Y. Zhang, Y. Hu, W. Wang, C. Zhang, Y. Gao, J. Hao, T. Lv, and C. Fan · 2022
Later among the works it cites.
Asynchronous actor-critic for multi-agent reinforcement learning
Y. Xiao, W. Tan, and C. Amato · 2022
Later among the works it cites.
Ldsa: Learning dynamic subtask assignment in cooperative multi-agent reinforcement learning
M. Yang, J. Zhao, X. Hu, W. Zhou, J. Zhu, and H. Li · 2022
Later among the works it cites.
Mcmarl: Parameterizing value function via mixture of categorical distributions for multi-agent reinforcement learning, 2022
J. Zhao, M. Yang, Y. Zhao, X. Hu, W. Zhou, J. Zhu, and H. Li · 2022
Later among the works it cites.
Who leads and who follows in strategic classification?, 2022
T. Zrnic, E. Mazumdar, S. S. Sastry, and M. I. Jordan · 2022
Later among the works it cites.
Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions, 2023
Y. Chebotar, Q. Vuong, A. Irpan, K. Hausman, F. Xia, Y. Lu, A. Kumar, T. Yu, A. Herzog, K. Pertsch, K. Gopalakrishnan, J. Ibarz, O. Nachum, S. Sontakke, G. Salazar, H. T. Tran, J. Peralta, C. Tan, D. Manjunath, J. Singht, B. Zitkovich, T. Jackson, K. Rao, C. Finn, and S. Levine · 2023
Closest in time.
Open x-embodiment: Robotic learning datasets and rt-x models, 2023
E. Collaboration, A. Padalkar, A. Pooley, A. Mandlekar, A. Jain, A. Tung, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Garg, A. Brohan, A. Raffin, A. Wahid, B. Burgess-Limerick, B. Kim, B. Schölkopf, B. Ichter, C. Lu, C. Xu, C. Finn, C. Xu, C. Chi, C. Huang, C. Chan, C. Pan, C. Fu, C. Devin, D. Driess, D. Pathak, D. Shah, D. Büchler, D. Kalashnikov, D. Sadigh, E. Johns, F. Ceola, F. Xia, F. Stulp, G. Zhou, G. S. Sukhatme, G. Salhotra, G. Yan, G. Schiavi, G. Kahn, H. Su, H.-S. Fang, H. Shi, H. B. Amor, H. I. Christensen, H. Furuta, H. Walke, H. Fang, I. Mordatch, I. Radosavovic, I. Leal, J. Liang, J. Abou-Chakra, J. Kim, J. Peters, J. Schneider, J. Hsu, J. Bohg, J. Bingham, J. Wu, J. Wu, J. Luo, J. Gu, J. Tan, J. Oh, J. Malik, J. Booher, J. Tompson, J. Yang, J. J. Lim, J. Silvério, J. Han, K. Rao, K. Pertsch, K. Hausman, K. Go, K. Gopalakrishnan, K. Goldberg, K. Byrne, K. Oslund, K. Kawaharazuka, K. Zhang, K. Rana, K. Srinivasan, L. Y. Chen, L. Pinto, L. Fei-Fei, L. Tan, L. Ott, L. Lee, M. Tomizuka, M. Spero, M. Du, M. Ahn, M. Zhang, M. Ding, M. K. Srirama, M. Sharma, M. J. Kim, N. Kanazawa, N. Hansen, N. Heess, N. J. Joshi, N. Suenderhauf, N. D. Palo, N. M. M. Shafiullah, O. Mees, O. Kroemer, P. R. Sanketi, P. Wohlhart, P. Xu, P. Sermanet, P. Sundaresan, Q. Vuong, R. Rafailov, R. Tian, R. Doshi, R. Martín-Martín, R. Mendonca, R. Shah, R. Hoque, R. Julian, S. Bustamante, S. Kirmani, S. Levine, S. Moore, S. Bahl, S. Dass, S. Sonawani, S. Song, S. Xu, S. Haldar, S. Adebola, S. Guist, S. Nasiriany, S. Schaal, S. Welker, S. Tian, S. Dasari, S. Belkhale, T. Osa, T. Harada, T. Matsushima, T. Xiao, T. Yu, T. Ding, T. Davchev, T. Z. Zhao, T. Armstrong, T. Darrell, V. Jain, V. Vanhoucke, W. Zhan, W. Zhou, W. Burgard, X. Chen, X. Wang, X. Zhu, X. Li, Y. Lu, Y. Chebotar, Y. Zhou, Y. Zhu, Y. Xu, Y. Wang, Y. Bisk, Y. Cho, Y. Lee, Y. Cui, Y.-H. Wu, Y. Tang, Y. Zhu, Y. Li, Y. Iwasawa, Y. Matsuo, Z. Xu, and Z. J. Cui · 2023
Closest in time.
Reinforcement learning from passive data via latent intentions, 2023
D. Ghosh, C. Bhateja, and S. Levine · 2023
Closest in time.
Safe multi-agent reinforcement learning for multi-robot control
S. Gu, J. G. Kuba, Y. Chen, Y. Du, L. Yang, A. Knoll, and Y. Yang · 2023
Closest in time.
Efficient human-ai coordination via preparatory language-based convention
C. Guan, L. Zhang, C. Fan, Y. Li, F. Chen, L. Li, Y. Tian, L. Yuan, and Y. Yu · 2023
Closest in time.
Explainable action advising for multi-agent reinforcement learning
Y. Guo, J. Campbell, S. Stepputtis, R. Li, D. Hughes, F. Fang, and K. Sycara · 2023
Closest in time.
Mastering diverse domains through world models, 2023
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap · 2023
Closest in time.
Exploration in deep reinforcement learning: From single-agent to multiagent domain
J. Hao, T. Yang, H. Tang, C. Bai, J. Liu, Z. Meng, P. Liu, and Z. Wang · 2023
Closest in time.
Decentralized multi-agent filtering, 2023
D. Huh and P. Mohapatra · 2023
Closest in time.
Optimal algorithms for decentralized stochastic variational inequalities, 2023
D. Kovalev, A. Beznosikov, A. Sadiev, M. Persiianov, P. Richtárik, and A. Gasnikov · 2023
Closest in time.
Gigastep - one billion steps per second multi-agent reinforcement learning
M. Lechner, L. Yin, T. Seyde, T.-H. Wang, W. Xiao, R. Hasani, J. Rountree, and D. Rus · 2023
Closest in time.
Lazy agents: a new perspective on solving sparse reward problem in multi-agent reinforcement learning
B. Liu, Z. Pu, Y. Pan, J. Yi, Y. Liang, and D. Zhang · 2023
Closest in time.
Towards few-shot coordination: Revisiting ad-hoc teamplay challenge in the game of hanabi
H. Nekoei, X. Zhao, J. Rajendran, M. Liu, and S. Chandar · 2023
Closest in time.
A modern introduction to online learning, 2023
F. Orabona · 2023
Closest in time.
Attention-based recurrence for multi-agent reinforcement learning under stochastic partial observability, 2023
T. Phan, F. Ritz, P. Altmann, M. Zorn, J. Nüßlein, M. Kölle, T. Gabor, and C. Linnhoff-Popien · 2023
Closest in time.
Jaxmarl: Multi-agent rl environments in jax
A. Rutherford, B. Ellis, M. Gallici, J. Cook, A. Lupu, G. Ingvarsson, T. Willi, A. Khan, C. S. de Witt, A. Souly, S. Bandyopadhyay, M. Samvelyan, M. Jiang, R. T. Lange, S. Whiteson, B. Lacerda, N. Hawes, T. Rocktaschel, C. Lu, and J. N. Foerster · 2023
Closest in time.
Learning from good trajectories in offline multi-agent reinforcement learning
Q. Tian, K. Kuang, F. Liu, and B. Wang · 2023
Closest in time.
Bridgedata v2: A dataset for robot learning at scale
H. Walke, K. Black, A. Lee, M. J. Kim, M. Du, C. Zheng, T. Zhao, P. Hansen-Estruch, Q. Vuong, A. He, V. Myers, K. Fang, C. Finn, and S. Levine · 2023
Closest in time.
Learning zero-shot cooperation with humans, assuming humans are biased
C. Yu, J. Gao, W. Liu, B. Xu, H. Tang, J. Yang, Y. Wang, and Y. Wu · 2023
Closest in time.
Multi-agent reinforcement learning: Foundations and modern approaches
S. V. Albrecht, F. Christianos, and L. Schäfer · 2024
Closest in time.
Policy space response oracles: A survey, 2024
A. Bighashdel, Y. Wang, S. McAleer, R. Savani, and F. A. Oliehoek · 2024
Closest in time.
A survey of multi-agent deep reinforcement learning with communication
C. Zhu, M. Dastani, and S. Wang · 2024
Closest in time.