Fetching the paper…
Reading the bibliography…
Generating agents that can achieve zero-shot coordination (ZSC) with unseen partners is a new challenge in cooperative multi-agent reinforcement learning (MARL).
G. Tesauro, “Td-gammon, a self-teaching backgammon program, achieves master-level play,” Neural Computation , vol. 6, no. 2, pp. 215–219, 1994
1994
Earlier work this paper cites.
T. Bäck, Evolutionary algorithms in theory and practice: Evolution Strategies, evolutionary programming, genetic algorithms . Oxford University Press, 1996
1996
Earlier work this paper cites.
M. A. Potter and K. A. D. Jong, “Cooperative coevolution: An architecture for evolving coadapted subcomponents,” Evolutionary Computation , vol. 8, no. 1, pp. 1–29, 2000
2000
Earlier work this paper cites.
S. Kapetanakis and D. Kudenko, “Reinforcement learning of coordination in heterogeneous cooperative multi-agent systems,” in Proceedings of the 3rd International Joint Conference on Autonomous Agents and Multiagent Systems (AAMAS) , New York, NY, 2004, pp. 1258–1259
2004
Earlier work this paper cites.
J. Demšar, “Statistical comparisons of classifiers over multiple data sets,” Journal of Machine Learning Research , vol. 7, pp. 1–30, 2006
2006
Earlier work this paper cites.
L. Van der Maaten and G. E. Hinton, “Visualizing data using t-sne,” Journal of machine learning research , vol. 9, no. 11, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
P. Stone, G. A. Kaminka, S. Kraus, and J. S. Rosenschein, “Ad hoc autonomous agent teams: Collaboration without pre-coordination,” in Proceedings of the 24th AAAI Conference on Artificial Intelligence (AAAI) , Atlanta, GA, 2010, p. 1504–1509
2010
Earlier work this paper cites.
J. Lehman and K. O. Stanley, “Abandoning objectives: Evolution through the search for novelty alone,” Evolutionary Computation , vol. 19, no. 2, pp. 189–223, 2011
2011
Earlier work this paper cites.
A. Cully, J. Clune, D. Tarapore, and J.-B. Mouret, “Robots that can adapt like animals,” Nature , vol. 521, no. 7553, pp. 503–507, 2015
2015
Earlier work this paper cites.
F. A. Oliehoek and C. Amato, A Concise Introduction to Decentralized POMDPs . Springer, 2016
2016
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van D. Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al. , “Mastering the game of go with deep neural networks and tree search,” Nature , vol. 529, no. 7587, pp. 484–489, 2016
2016
Earlier work this paper cites.
J. C. Brant and K. O. Stanley, “Minimal criterion coevolution: A new approach to open-ended search,” in Proceedings of the 19th ACM Genetic and Evolutionary Computation Conference (GECCO) , Berlin, Germany, 2017, pp. 67–74
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
E. Conti, V. Madhavan, F. P. Such, J. Lehman, K. O. Stanley, and J. Clune, “Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents,” in Advances in Neural Information Processing Systems 31 (NeurIPS) , Montréal, Canada, 2018, pp. 5032–5043
2018
Earlier work this paper cites.
A. Cully and Y. Demiris, “Quality and diversity optimization: A unifying modular framework,” IEEE Transactions on Evolutionary Computation , vol. 22, no. 2, pp. 245–259, 2018
2018
Earlier work this paper cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in Proceedings of the 35th International Conference on Machine Learning (ICML) , Stockholm, Sweden, 2018, pp. 1861–1870
2018
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction . MIT Press, 2018
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Carroll, R. Shah, M. K. Ho, T. Griffiths, S. Seshia, P. Abbeel, and A. Dragan, “On the utility of learning about humans for human-ai coordination,” in Advances in Neural Information Processing Systems 32 (NeurIPS) , Vancouver, Canada, 2019, pp. 5175–5186
2019
Cited alongside, same era.
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine, “Diversity is all you need: Learning skills without a reward function,” in Proceedings of the 6th International Conference on Learning Representations (ICLR) , 2019
2019
Cited alongside, same era.
X. Ma, X. Li, Q. Zhang, K. Tang, Z. Liang, W. Xie, and Z. Zhu, “A survey on cooperative co-evolutionary algorithms,” IEEE Transactions on Evolutionary Computation , vol. 23, no. 3, pp. 421–441, 2019
2019
Cited alongside, same era.
R. Wang, J. Lehman, J. Clune, and K. O. Stanley, “POET: Open-ended coevolution of environments and their optimized solutions,” in Proceedings of the 21st ACM Genetic and Evolutionary Computation Conference (GECCO) , Prague, Czech Republic, 2019, pp. 142–151
W. Mondal, M. Agarwal, V. Aggarwal, and S. V. Ukkusuri, “On the approximation of cooperative heterogeneous multi-agent reinforcement learning (MARL) using mean field control (MFC),” Journal of Machine Learning Research , vol. 23, no. 129, pp. 1–46, 2022
2022
Closest in time.
2022
Closest in time.
A. Salehi, A. Coninx, and S. Doncieux, “Few-shot quality-diversity optimization,” IEEE Robotics and Automation Letters , vol. 7, no. 2, pp. 4424–4431, 2022
2022
Closest in time.
C. Wang, C. Pérez-D’Arpino, D. Xu, F.-F. Li, K. Liu, and S. Savarese, “Co-gail: Learning diverse strategies for human-robot collaboration,” in Proceedings of the 6th Conference on Robot Learning (CoRL) , Auckland, NZ, 2022, pp. 1279–1290
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
R. Zhao, J. Song, H. Hu, Y. Gao, Y. Wu, Z. Sun, and Y. Wei, “Maximum entropy population based training for zero-shot human-ai coordination,” in Proceedings of the 37th AAAI Conference on Artificial Intelligence (AAAI) , Washington, DC, 2023, pp. 6145–6153
2019
Cited alongside, same era.
H. Hu, A. Lerer, A. Peysakhovich, and J. N. Foerster, ““other-play” for zero-shot coordination,” in Proceedings of the 37th International Conference on Machine Learning (ICML) , Vienna, Austria, 2020, pp. 4399–4410
2020
Cited alongside, same era.
S. Kumar, A. Kumar, S. Levine, and C. Finn, “One solution is not all you need: Few-shot extrapolation via structured MaxEnt RL,” in Advances in Neural Information Processing Systems 34 (NeurIPS) , Vancouver, Canada, 2020, pp. 8198–8210
2020
Cited alongside, same era.
J. Parker-Holder, L. Metz, C. Resnick, H. Hu, A. Lerer, A. Letcher, A. Peysakhovich, A. Pacchiano, and J. Foerster, “Ridge rider: Finding diverse solutions by following eigenvectors of the hessian,” in Advances in Neural Information Processing Systems 33 (NeurIPS) , Vancouver, Canada, 2020, pp. 753–765
2020
Cited alongside, same era.
J. Parker-Holder, A. Pacchiano, K. M. Choromanski, and S. J. Roberts, “Effective diversity in population based reinforcement learning,” in Advances in Neural Information Processing Systems 33 (NeurIPS) , Vancouver, Canada, 2020, pp. 18 050–18 062
2020
Cited alongside, same era.
R. Portelas, C. Colas, L. Weng, K. Hofmann, and P. Oudeyer, “Automatic curriculum learning for deep RL: A short survey,” in Proceedings of the 29th International Joint Conference on Artificial Intelligence (IJCAI) , Yokohama, Japan, 2020, pp. 4819–4825
2020
Cited alongside, same era.
R. Wang, J. Lehman, A. Rawal, J. Zhi, Y. Li, J. Clune, and K. O. Stanley, “Enhanced POET: Open-ended reinforcement learning through unbounded invention of learning challenges and their solutions,” in Proceedings of the 37th International Conference on Machine Learning (ICML) , Vienna, Austria, 2020, pp. 9940–9951
2020
Cited alongside, same era.
K. Chatzilygeroudis, A. Cully, V. Vassiliades, and J.-B. Mouret, “Quality-Diversity optimization: A novel branch of stochastic optimization,” in Black Box Optimization, Machine Learning, and No-Free Lunch Theorems . Springer, 2021, pp. 109–135
2021
Cited alongside, same era.
Y. Wang, K. Xue, and C. Qian, “Evolutionary diversity optimization with clustering-based selection for reinforcement learning,” in Proceedings of the 10th International Conference on Learning Representations (ICLR) , 2022
2022
Closest in time.
W. Xue, W. Qiu, B. An, Z. Rabinovich, S. Obraztsova, and C. K. Yeo, “Mis-spoke or mis-lead: Achieving robustness in multi-agent communicative reinforcement learning,” in Proceedings of the 21st International Conference on Autonomous Agents and Multiagent Systems , 2022, pp. 1418–1426
2022
Closest in time.
R. Charakorn, P. Manoonpong, and N. Dilokthanakul, “Generating diverse cooperative agents by learning incompatible policies,” in The 11th International Conference on Learning Representations (ICLR) , Kigali, Rwanda, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Y. Li, S. Zhang, J. Sun, Y. Du, Y. Wen, X. Wang, and W. Pan, “Cooperative open-ended learning framework for zero-shot coordination,” in Proceedings of the 40th International Conference on Machine Learning (ICML) , Honolulu, HI, 2023
2023
Closest in time.
Y. Loo, C. Gong, and M. Meghjani, “A hierarchical approach to population training for human-ai collaboration,” in Proceedings of the 32nd International Joint Conference on Artificial Intelligence (IJCAI) , Macao, SAR, China, 2023, pp. 3011–3019
2023
Closest in time.
X. Lou, J. Guo, J. Zhang, J. Wang, K. Huang, and Y. Du, “PECAN: leveraging policy ensemble for context-aware zero-shot human-ai coordination,” in Proceedings of the 22nd International Conference on Autonomous Agents and Multiagent Systems (AAMAS) , London, United Kingdom, 2023, pp. 679–688
2023
Closest in time.
C. Yu, J. Gao, W. Liu, B. Xu, H. Tang, J. Yang, Y. Wang, and Y. Wu, “Learning zero-shot cooperation with humans, assuming humans are biased,” in The 11th International Conference on Learning Representations (ICLR) , Kigali, Rwanda, 2023
2023
Closest in time.
2023
Closest in time.
L. Yuan, Z. Zhang, K. Xue, H. Yin, F. Chen, C. Guan, L. Li, C. Qian, and Y. Yu, “Robust multi-agent coordination via evolutionary generation of auxiliary adversarial attackers,” in Proceedings of the 37th AAAI Conference on Artificial Intelligence (AAAI) , Washington, DC, 2023
2023
Closest in time.
H. Hao, X. Zhang, and A. Zhou, “Enhancing SAEAs with unevaluated solutions: a case study of relation model for expensive optimization,” Science China Information Sciences , vol. 67, no. 2, p. 120103, 2024
2024
Closest in time.
R. Jiao, B. H. Nguyen, B. Xue, and M. Zhang, “A survey on evolutionary multiobjective feature selection in classification: Approaches, applications, and challenges,” IEEE Transactions on Evolutionary Computation , vol. 28, no. 4, pp. 1156–1176, 2024
2024
Closest in time.
J. Liang, Y. Zhang, K. Chen, B. Qu, K. Yu, C. Yue, and P. N. Suganthan, “An evolutionary multiobjective method based on dominance and decomposition for feature selection in classification,” Science China Information Sciences , vol. 67, no. 2, p. 120101, 2024
2024
Closest in time.
C. Qian, K. Xue, and R. Wang, “Quality-diversity algorithms can provably be helpful for optimization,” in Proceedings of the 31st International Joint Conference on Artificial Intelligence (IJCAI) , 2024
2024
Closest in time.
L. Yuan, F. Chen, Z. Zhang, and Y. Yu, “Communication-robust multi-agent learning by adaptable auxiliary multi-agent adversary generation,” Frontiers of Computer Science , vol. 18, no. 6, p. 186331, 2024
2024
Closest in time.