Fetching the paper…
Reading the bibliography…
Sample efficiency remains a key challenge in multi-agent reinforcement learning (MARL).
Stochastic games*
L. S. Shapley · 1953
Earlier work this paper cites.
Games with incomplete information played by “bayesian” players, i–iii part i. the basic model
J. C. Harsanyi · 1967
Earlier work this paper cites.
Improved baselines with momentum contrastive learning. arxiv 2020
X. Chen, H. Fan, R. Girshick, and K. He · 2003
Earlier work this paper cites.
Equivalence notions and model minimization in markov decision processes
R. Givan, T. Dean, and M. Greig · 2003
Earlier work this paper cites.
Multi-agent reinforcement learning: An overview
L. Buşoniu, R. Babuška, and B. De Schutter · 2010
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning, 2016
Y. Gal and Z. Ghahramani · 2016
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. I. Wu, A. Tamar, J. Harb, O. Pieter Abbeel, and I. Mordatch · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
How many random seeds? statistical power analysis in deep reinforcement learning experiments
C. Colas, O. Sigaud, and P.-Y. Oudeyer · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Hoof, and D. Meger · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Recurrent experience replay in distributed reinforcement learning
S. Kapturowski, G. Ostrovski, J. Quan, R. Munos, and W. Dabney · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2019
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Earlier work this paper cites.
Bootstrap your own latent-a new approach to self-supervised learning
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al · 2020
Earlier work this paper cites.
Bootstrap latent-predictive representations for multitask reinforcement learning
Z. D. Guo, B. A. Pires, B. Piot, J.-B. Grill, F. Altché, R. Munos, and M. G. Azar · 2020
Earlier work this paper cites.
Curl: Contrastive unsupervised representations for reinforcement learning
M. Laskin, A. Srinivas, and P. Abbeel · 2020
Cited alongside, same era.
Independent Learning Approaches: Overcoming Multi-Agent Learning Pathologies In Team-Games
G. Palmer · 2020
Cited alongside, same era.
Data-efficient reinforcement learning with self-predictive representations
M. Schwarzer, A. Anand, R. Goel, R. D. Hjelm, A. Courville, and P. Bachman · 2020
Cited alongside, same era.
Sample-efficient optimization in the latent space of deep generative models via weighted retraining
A. Tripp, E. Daxberger, and J. M. Hernández-Lobato · 2020
Cited alongside, same era.
Phasic policy gradient
K. W. Cobbe, J. Hilton, O. Klimov, and J. Schulman · 2021
Cited alongside, same era.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
Offline multi-agent reinforcement learning with knowledge distillation
W.-C. Tseng, T.-H. J. Wang, Y.-C. Lin, and P. Isola · 2022
Later among the works it cites.
Multi-agent reinforcement learning is a sequence modeling problem
M. Wen, J. Kuba, R. Lin, W. Zhang, Y. Wen, J. Wang, and Y. Yang · 2022
Later among the works it cites.
The surprising effectiveness of ppo in cooperative multi-agent games
C. Yu, A. Velu, E. Vinitsky, J. Gao, Y. Wang, A. Bayen, and Y. Wu · 2022
Later among the works it cites.
BenchMARL: Benchmarking Multi-Agent Reinforcement Learning
M. Bettini, A. Prorok, and V. Moens · 2023
Later among the works it cites.
The complexity of markov equilibrium in stochastic games
C. Daskalakis, N. Golowich, and K. Zhang · 2023
Later among the works it cites.
Joint-predictive representations for multi-agent reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Cited alongside, same era.
Trust region policy optimisation in multi-agent reinforcement learning
J. G. Kuba, R. Chen, M. Wen, Y. Wen, F. Sun, J. Wang, and Y. Yang · 2021
Cited alongside, same era.
Pretraining representations for data-efficient reinforcement learning
M. Schwarzer, N. Rajkumar, M. Noukhovitch, A. Anand, L. Charlin, R. D. Hjelm, P. Bachman, and A. C. Courville · 2021
Cited alongside, same era.
Agent-centric representations for multi-agent reinforcement learning
W. Shang, L. Espeholt, A. Raichuk, and T. Salimans · 2021
Cited alongside, same era.
Decoupling representation learning from reinforcement learning
A. Stooke, K. Lee, P. Abbeel, and M. Laskin · 2021
Cited alongside, same era.
Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning
Y. Yang, X. Ma, C. Li, Z. Zheng, Q. Zhang, G. Huang, J. Yang, and Q. Zhao · 2021
Cited alongside, same era.
Vmas: A vectorized multi-agent simulator for collective robot learning
M. Bettini, R. Kortvelesy, J. Blumenkamp, and A. Prorok · 2022
Cited alongside, same era.
M. Feng, W. Zhou, Y. Yang, and H. Li · 2023
Later among the works it cites.
Mastering diverse domains through world models
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap · 2023
Later among the works it cites.
Isaacteams: Extending gpu-based physics simulator for multi-agent learning, 2023
D. Huh and P. Mohapatra · 2023
Later among the works it cites.
Sample-efficient multi-agent reinforcement learning with masked reconstruction
J. I. Kim, Y. J. Lee, J. Heo, J. Park, J. Kim, S. R. Lim, J. Jeong, and S. B. Kim · 2023
Later among the works it cites.
Two heads are better than one: A simple exploration framework for efficient multi-agent reinforcement learning
J. Li, K. Kuang, B. Wang, X. Li, F. Wu, J. Xiao, and L. Chen · 2023
Later among the works it cites.
Efficient multi-agent reinforcement learning by planning
Q. Liu, J. Ye, X. Ma, J. Yang, B. Liang, and C. Zhang · 2023
Later among the works it cites.
Ma2cl: Masked attentive contrastive learning for multi-agent reinforcement learning
H. Song, M. Feng, W. Zhou, and H. Li · 2023
Later among the works it cites.
Understanding self-predictive learning for reinforcement learning
Y. Tang, Z. D. Guo, P. H. Richemond, B. A. Pires, Y. Chandak, R. Munos, M. Rowland, M. G. Azar, C. Le Lan, C. Lyle, et al · 2023
Later among the works it cites.
Deep latent state space models for time-series generation
L. Zhou, M. Poli, W. Xu, S. Massaroli, and S. Ermon · 2023
Later among the works it cites.
Attention-guided contrastive role representations for multi-agent reinforcement learning
Z. Hu, Z. Zhang, H. Li, C. Chen, H. Ding, and Z. Wang · 2024
Closest in time.
Multi-agent reinforcement learning: A comprehensive survey, 2024
D. Huh and P. Mohapatra · 2024
Closest in time.
Dreamsmooth: Improving model-based reinforcement learning via reward smoothing
V. Lee, P. Abbeel, and Y. Lee · 2024
Closest in time.
Bridging state and history representations: Understanding self-predictive rl
T. Ni, B. Eysenbach, E. Seyedsalehi, M. Ma, C. Gehring, A. Mahajan, and P.-L. Bacon · 2024
Closest in time.