Fetching the paper…
Reading the bibliography…
Multi-agent reinforcement learning (MARL) requires agents to explore within a vast joint action space to find joint actions that lead to coordination.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson. 1933 · 1933
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents. In International Conference on Machine Learning
Ming Tan. 1993 · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman. 1994 · 1994
Earlier work this paper cites.
Bayesian Q-learning
Richard Dearden, Nir Friedman, and Stuart Russell. 1998 · 1998
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer. 2002 · 2002
Earlier work this paper cites.
The complexity of decentralized control of Markov decision processes
Daniel S. Bernstein, Robert Givan, Neil Immerman, and Shlomo Zilberstein. 2002 · 2002
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games. In AAAI Conference on Artificial Intelligence
Eric A. Hansen, Daniel S. Bernstein, and Shlomo Zilberstein. 2004 · 2004
Earlier work this paper cites.
A modern Bayesian look at the multi-armed bandit
Steven L. Scott. 2010 · 2010
Earlier work this paper cites.
An empirical evaluation of Thompson sampling. In Advances in Neural Information Processing Systems
Olivier Chapelle and Lihong Li. 2011 · 2011
Earlier work this paper cites.
A Game-Theoretic Model and Best-Response Learning Method for Ad Hoc Coordination in Multiagent Systems. In International Conference on Autonomous Agents and Multi-Agent Systems
Stefano V. Albrecht and Subramanian Ramamoorthy. 2013 · 2013
Earlier work this paper cites.
(More) efficient reinforcement learning via posterior sampling. In Advances in Neural Information Processing Systems
Ian Osband, Daniel Russo, and Benjamin van Roy. 2013 · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Conference on Empirical Methods in Natural Language Processing
Kyunghyun Cho, Bart Van Merriënboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. 2015 · 2015
Earlier work this paper cites.
A Concise Introduction to Decentralized POMDPs . Vol. 1
Frans A. Oliehoek and Christopher Amato. 2016 · 2016
Earlier work this paper cites.
Averaged-DQN: Variance reduction and stabilization for deep reinforcement learning. In International Conference on Machine Learning
Oron Anschel, Nir Baram, and Nahum Shimkin. 2017 · 2017
Earlier work this paper cites.
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. In Advances in Neural Information Processing Systems
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Earlier work this paper cites.
Why is posterior sampling better than optimism for reinforcement learning?. In International Conference on Machine Learning
Ian Osband and Benjamin van Roy. 2017 · 2017
Cited alongside, same era.
Emergence of grounded compositional language in multi-agent populations. In AAAI Conference on Artificial Intelligence
Igor Mordatch and Pieter Abbeel. 2018 · 2018
Cited alongside, same era.
Value-Decomposition networks for cooperative multi-agent learning. In International Conference on Autonomous Agents and Multi-Agent Systems
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech M. Czarnecki, Vinícius Flores Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel. 2018 · 2018
Cited alongside, same era.
Better Exploration with Optimistic Actor Critic. In Advances in Neural Information Processing Systems
Kamil Ciosek, Quan Vuong, Robert Loftin, and Katja Hofmann. 2019 · 2019
Cited alongside, same era.
LIIR: Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning. In Advances in Neural Information Processing Systems
Scaling multi-agent reinforcement learning with selective parameter sharing. In International Conference on Machine Learning
Filippos Christianos, Georgios Papoudakis, Muhammad A. Rahman, and Stefano V. Albrecht. 2021 · 2021
Later among the works it cites.
SUNRISE: A simple unified framework for ensemble learning in deep reinforcement learning. In International Conference on Machine Learning
Kimin Lee, Michael Laskin, Aravind Srinivas, and Pieter Abbeel. 2021 · 2021
Later among the works it cites.
Celebrating diversity in shared multi-agent reinforcement learning. In Advances in Neural Information Processing Systems
Chenghao Li, Tonghan Wang, Chengjie Wu, Qianchuan Zhao, Jun Yang, and Chongjie Zhang. 2021 · 2021
Later among the works it cites.
Cooperative exploration for multi-agent deep reinforcement learning. In International Conference on Machine Learning
Iou-Jen Liu, Unnat Jain, Raymond A. Yeh, and Alexander Schwing. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yali Du, Lei Han, Meng Fang, Ji Liu, Tianhong Dai, and Dacheng Tao. 2019 · 2019
Cited alongside, same era.
Successor uncertainties: Exploration and uncertainty in temporal difference learning. In Advances in Neural Information Processing Systems
David Janz, Jiri Hron, Przemysław Mazur, Katja Hofmann, José Miguel Hernández-Lobato, and Sebastian Tschiatschek. 2019 · 2019
Cited alongside, same era.
MAVEN: Multi-Agent Variational Exploration. In Advances in Neural Information Processing Systems
Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, and Shimon Whiteson. 2019 · 2019
Cited alongside, same era.
Deep Exploration via Randomized Value Functions
Ian Osband, Benjamin van Roy, Daniel J. Russo, and Zheng Wen. 2019 · 2019
Cited alongside, same era.
Dealing with Non-Stationarity in Multi-Agent Deep Reinforcement Learning
Georgios Papoudakis, Filippos Christianos, Arrasy Rahman, and Stefano V. Albrecht. 2019 · 2019
Cited alongside, same era.
Measuring the reliability of reinforcement learning algorithms. In International Conference on Learning Representations
Stephanie C.Y. Chan, Samuel Fishman, John Canny, Anoop Korattikara, and Sergio Guadarrama. 2020 · 2020
Cited alongside, same era.
Shared Experience Actor-Critic for Multi-Agent Reinforcement Learning. In Advances in Neural Information Processing Systems
Filippos Christianos, Lukas Schäfer, and Stefano V. Albrecht. 2020 · 2020
Cited alongside, same era.
Hypermodels for exploration. In International Conference on Learning Representations
Vikranth Dwaracherla, Xiuyuan Lu, Morteza Ibrahimi, Ian Osband, Zheng Wen, and Benjamin van Roy. 2020 · 2020
Cited alongside, same era.
Georgios Papoudakis, Filippos Christianos, Lukas Schäfer, and Stefano V. Albrecht. 2021 · 2021
Later among the works it cites.
Episodic multi-agent reinforcement learning with curiosity-driven exploration. In Advances in Neural Information Processing Systems
Lulu Zheng, Jiarui Chen, Jianhao Wang, Jiamin He, Yujing Hu, Yingfeng Chen, Changjie Fan, Yang Gao, and Chongjie Zhang. 2021 · 2021
Later among the works it cites.
Model-based Lifelong Reinforcement Learning with Bayesian Exploration. In Advances in Neural Information Processing Systems
Haotian Fu, Shangqun Yu, Michael Littman, and George Konidaris. 2022 · 2022
Later among the works it cites.
Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep Networks. In International Conference on Machine Learning
Litian Liang, Yaosheng Xu, Stephen McAleer, Dailin Hu, Alexander Ihler, Pieter Abbeel, and Roy Fox. 2022 · 2022
Later among the works it cites.
LIGS: Learnable Intrinsic-Reward Generation Selection for Multi-Agent Learning. In International Conference on Learning Representations
David Henry Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez-Nieves, Oliver Slumbers, Feifei Tong, Yang Li, Jiangcheng Zhu, Yaodong Yang, and Jun Wang. 2022 · 2022
Later among the works it cites.
Decoupled Reinforcement Learning to Stabilise Intrinsically-Motivated Exploration. In International Conference on Autonomous Agents and Multiagent Systems
Lukas Schäfer, Filippos Christianos, Josiah P. Hanna, and Stefano V. Albrecht. 2022 · 2022
Later among the works it cites.
Efficient Model-based Multi-agent Reinforcement Learning via Optimistic Equilibrium Computation. In International Conference on Machine Learning
Pier Giuseppe Sessa, Maryam Kamgarpour, and Andreas Krause. 2022 · 2022
Later among the works it cites.
The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games. In Advances in Neural Information Processing Systems, Track on Datasets and Benchmarks
Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre Bayen, and Yi Wu. 2022 · 2022
Later among the works it cites.
Pareto Actor-Critic for Equilibrium Selection in Multi-Agent Reinforcement Learning
Filippos Christianos, Georgios Papoudakis, and Stefano V. Albrecht. 2023 · 2023
Closest in time.
Implicit Ensemble Training for Efficient and Robust Multiagent Reinforcement Learning
Macheng Shen and Jonathan P. How. 2023 · 2023
Closest in time.
Multi-Agent Reinforcement Learning: Foundations and Modern Approaches
Stefano V. Albrecht, Filippos Christianos, and Lukas Schäfer. 2024 · 2024
Closest in time.