Fetching the paper…
Reading the bibliography…
Self-play, a learning paradigm where agents iteratively refine their policies by interacting with historical or concurrent versions of themselves or other evolving agents, has shown remarkable success in solving complex non-cooperative multi-agent tasks.
Non-cooperative games
John F Nash et al · 1950
Earlier work this paper cites.
Iterative solution of games by fictitious play
George W Brown. 1951 · 1951
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley. 1953 · 1953
Earlier work this paper cites.
Some studies in machine learning using the game of checkers
Arthur L Samuel. 1959 · 1959
Earlier work this paper cites.
Subjectivity and correlation in randomized strategies
Robert J Aumann. 1974 · 1974
Earlier work this paper cites.
The rating of chessplayers: Past and present
Arpad E Elo and Sam Sloan. 1978 · 1978
Earlier work this paper cites.
Evolutionary stable strategies and game dynamics
Peter D Taylor and Leo B Jonker. 1978 · 1978
Earlier work this paper cites.
A world championship caliber checkers program
Jonathan Schaeffer, Joseph Culberson, Norman Treloar, Brent Knight, Paul Lu, and Duane Szafron. 1992 · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman. 1994 · 1994
Earlier work this paper cites.
On-line policy improvement using Monte-Carlo search
Gerald Tesauro and Gregory Galperin. 1996 · 1996
Earlier work this paper cites.
From simple features to sophisticated evaluation functions. In Computers and Games: First International Conference, CG’98 Tsukuba, Japan, November 11–12, 1998 Proceedings 1 . Springer, 126–145
Michael Buro. 1999 · 1998
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Yoav Freund and Robert E Schapire. 1999 · 1999
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Sergiu Hart and Andreu Mas-Colell. 2000 · 2000
Earlier work this paper cites.
Deep blue
Murray Campbell, A Joseph Hoane Jr, and Feng-hsiung Hsu. 2002 · 2002
Earlier work this paper cites.
World-championship-caliber Scrabble
Brian Sheppard. 2002 · 2002
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary. In Proceedings of the 20th International Conference on Machine Learning (ICML-03) . 536–543
H Brendan McMahan, Geoffrey J Gordon, and Avrim Blum. 2003 · 2003
Earlier work this paper cites.
Monte-carlo go developments
Bruno Bouzy and Bernard Helmstetter. 2004 · 2004
Earlier work this paper cites.
Efficient selectivity and backup operators in Monte-Carlo tree search. In International conference on computers and games . Springer, 72–83
Rémi Coulom. 2006 · 2006
Earlier work this paper cites.
TrueSkill™: a Bayesian skill rating system
Ralf Herbrich, Tom Minka, and Thore Graepel. 2006 · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning. In European conference on machine learning . Springer, 282–293
Levente Kocsis and Csaba Szepesvári. 2006 · 2006
Earlier work this paper cites.
Computing “elo ratings” of move patterns in the game of go
Rémi Coulom. 2007 · 2007
Earlier work this paper cites.
Combining online and offline knowledge in UCT. In Proceedings of the 24th international conference on Machine learning . 273–280
Sylvain Gelly and David Silver. 2007 · 2007
Earlier work this paper cites.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione. 2007 · 2007
Earlier work this paper cites.
The complexity of computing a Nash equilibrium
Constantinos Daskalakis, Paul W Goldberg, and Christos H Papadimitriou. 2009 · 2009
Earlier work this paper cites.
Monte Carlo sampling for regret minimization in extensive games
Marc Lanctot, Kevin Waugh, Martin Zinkevich, and Michael Bowling. 2009 · 2009
Earlier work this paper cites.
Fuego—an open-source framework for board games and Go engine based on Monte Carlo tree search
Markus Enzenberger, Martin Müller, Broderick Arneson, and Richard Segal. 2010 · 2010
Earlier work this paper cites.
Pachi: State of the art open source Go program
Petr Baudiš and Jean-loup Gailly. 2011 · 2011
Earlier work this paper cites.
Determinantal point processes for machine learning
Alex Kulesza, Ben Taskar, et al · 2012
Earlier work this paper cites.
On the complexity of approximating a Nash equilibrium
Constantinos Daskalakis. 2013 · 2013
Earlier work this paper cites.
The complexity of the homotopy method, equilibrium selection, and Lemke-Howson solutions
Paul W Goldberg, Christos H Papadimitriou, and Rahul Savani. 2013 · 2013
Earlier work this paper cites.
Measuring the size of large no-limit poker games
Michael Johanson. 2013 · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013 · 2013
Earlier work this paper cites.
Solving imperfect information games using decomposition. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 28
Neil Burch, Michael Johanson, and Michael Bowling. 2014 · 2014
Earlier work this paper cites.
Solving large imperfect information games using CFR+
Oskari Tammelin. 2014 · 2014
Earlier work this paper cites.
Heads-up limit hold’em poker is solved
Michael Bowling, Neil Burch, Michael Johanson, and Oskari Tammelin. 2015 · 2015
Earlier work this paper cites.
Regret-based pruning in extensive-form games
Noam Brown and Tuomas Sandholm. 2015 · 2015
Earlier work this paper cites.
Fictitious self-play in extensive-form games. In International conference on machine learning . PMLR, 805–813
Johannes Heinrich, Marc Lanctot, and David Silver. 2015 · 2015
Earlier work this paper cites.
Building a computer Mahjong player based on Monte Carlo simulation and opponent models. In 2015 IEEE Conference on Computational Intelligence and Games (CIG) . IEEE, 275–283
Naoki Mizukami and Yoshimasa Tsuruoka. 2015 · 2015
Earlier work this paper cites.
Solving games with functional regret estimation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 29
Kevin Waugh, Dustin Morrill, James Bagnell, and Michael Bowling. 2015 · 2015
Earlier work this paper cites.
Strategy-based warm starting for regret minimization in games. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 30
Noam Brown and Tuomas Sandholm. 2016 · 2016
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando De Freitas, and Shimon Whiteson. 2016 · 2016
Earlier work this paper cites.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver. 2016 · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon. 2016 · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Poker-CNN: A pattern learning strategy for making draws and bets in poker games using convolutional networks. In Proceedings of the aaai conference on artificial intelligence , Vol. 30
Nikolai Yakovenko, Liangliang Cao, Colin Raffel, and James Fan. 2016 · 2016
Earlier work this paper cites.
Dynamic thresholding and pruning for regret minimization. In Proceedings of the AAAI conference on artificial intelligence , Vol. 31
Noam Brown, Christian Kroer, and Tuomas Sandholm. 2017 · 2017
Earlier work this paper cites.
Reduced space and faster convergence in imperfect-information games via pruning. In International conference on machine learning . PMLR, 596–604
Noam Brown and Tuomas Sandholm. 2017 · 2017
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolat, David Silver, and Thore Graepel. 2017 · 2017
Earlier work this paper cites.
Deal or No Deal? End-to-End Learning of Negotiation Dialogues. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . 2443–2453
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Cited alongside, same era.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisỳ, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling. 2017 · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Emergent Complexity via Multi-Agent Competition. In International Conference on Learning Representations
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch. 2018 · 2018
Cited alongside, same era.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm. 2018 · 2018
Towards unifying behavioral and response diversity for open-ended learning in zero-sum games
Xiangyu Liu, Hangtian Jia, Ying Wen, Yujing Hu, Yingfeng Chen, Changjie Fan, Zhipeng Hu, and Yaodong Yang. 2021 · 2021
Later among the works it cites.
Multi-agent training beyond zero-sum with correlated equilibrium meta-solvers. In International Conference on Machine Learning . PMLR, 7480–7491
Luke Marris, Paul Muller, Marc Lanctot, Karl Tuyls, and Thore Graepel. 2021 · 2021
Later among the works it cites.
XDO: A double oracle algorithm for extensive-form games
Stephen McAleer, John B Lanier, Kevin A Wang, Pierre Baldi, and Roy Fox. 2021 · 2021
Later among the works it cites.
Modelling behavioural diversity for learning in open-ended games. In International conference on machine learning . PMLR, 8514–8524
Nicolas Perez-Nieves, Yaodong Yang, Oliver Slumbers, David H Mguni, Ying Wen, and Jun Wang. 2021 · 2021
Later among the works it cites.
From poincaré recurrence to convergence in imperfect information games: Finding equilibrium via regularization. In International Conference on Machine Learning . PMLR, 8525–8535
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Counterfactual multi-agent policy gradients. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 32
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. 2018 · 2018
Cited alongside, same era.
Regret minimization for partially observable deep reinforcement learning. In International conference on machine learning . PMLR, 2342–2351
Peter Jin, Kurt Keutzer, and Sergey Levine. 2018 · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning. In International Conference on Machine Learning . PMLR, 4295–4304
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. 2018 · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Cited alongside, same era.
Open-ended learning in symmetric zero-sum games. In International Conference on Machine Learning . PMLR, 434–443
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech Czarnecki, Julien Perolat, Max Jaderberg, and Thore Graepel. 2019 · 2019
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemyslaw Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Cited alongside, same era.
Julien Perolat, Remi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei, Mark Rowland, Pedro Ortega, Neil Burch, Thomas Anthony, David Balduzzi, Bart De Vylder, et al · 2021
Later among the works it cites.
Collaborating with humans without human data
DJ Strouse, Kevin McKee, Matt Botvinick, Edward Hughes, and Richard Everett. 2021 · 2021
Later among the works it cites.
SCC: An efficient deep reinforcement learning agent mastering the game of StarCraft II. In International conference on machine learning . PMLR, 10905–10915
Xiangjun Wang, Junxiao Song, Penghui Qi, Peng Peng, Zhenkun Tang, Wei Zhang, Weimin Li, Xiongjun Pi, Jujie He, Chao Gao, et al · 2021
Later among the works it cites.
Douzero: Mastering doudizhu with self-play deep reinforcement learning. In international conference on machine learning . PMLR, 12333–12344
Daochen Zha, Jingru Xie, Wenye Ma, Sheng Zhang, Xiangru Lian, Xia Hu, and Ji Liu. 2021 · 2021
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Later among the works it cites.
Online double oracle
Le Cong Dinh, Stephen McAleer, Zheng Tian, Nicolas Perez-Nieves, Oliver Slumbers, David Henry Mguni, Jun Wang, Haitham Bou Ammar, and Yaodong Yang. 2022 · 2022
Later among the works it cites.
From motor control to team play in simulated humanoid football
Siqi Liu, Guy Lever, Zhe Wang, Josh Merel, SM Ali Eslami, Daniel Hennes, Wojciech M Czarnecki, Yuval Tassa, Shayegan Omidshafiei, Abbas Abdolmaleki, et al · 2022
Later among the works it cites.
Self-play psro: Toward optimal populations in two-player zero-sum games
Stephen McAleer, John Banister Lanier, Kevin Wang, Pierre Baldi, Roy Fox, and Tuomas Sandholm. 2022a · 2022
Later among the works it cites.
Anytime psro for two-player zero-sum games
Stephen McAleer, Kevin Wang, John Lanier, Marc Lanctot, Pierre Baldi, Tuomas Sandholm, and Roy Fox. 2022b · 2022
Later among the works it cites.
Human-level play in the game of Diplomacy by combining language models with strategic reasoning
Meta, Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, et al · 2022
Later among the works it cites.
Learning Equilibria in Mean-Field Games: Introducing Mean-Field PSRO. In Proceedings of the 21st International Conference on Autonomous Agents and Multiagent Systems . 926–934
Paul Muller, Mark Rowland, Romuald Elie, Georgios Piliouras, Julien Perolat, Mathieu Lauriere, Raphael Marinier, Olivier Pietquin, and Karl Tuyls. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Mastering the game of Stratego with model-free multiagent reinforcement learning
Julien Perolat, Bart De Vylder, Daniel Hennes, Eugene Tarassov, Florian Strub, Vincent de Boer, Paul Muller, Jerome T Connor, Neil Burch, Thomas Anthony, et al · 2022
Later among the works it cites.
Perfectdou: Dominating doudizhu with perfect information distillation
Guan Yang, Minghuan Liu, Weijun Hong, Weinan Zhang, Fei Fang, Guangjun Zeng, and Yue Lin. 2022 · 2022
Later among the works it cites.
The surprising effectiveness of ppo in cooperative multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. 2022 · 2022
Later among the works it cites.
Douzero+: Improving doudizhu ai by opponent modeling and coach-guided learning. In 2022 IEEE Conference on Games (CoG) . IEEE, 127–134
Youpeng Zhao, Jian Zhao, Xunhan Hu, Wengang Zhou, and Houqiang Li. 2022b · 2022
Later among the works it cites.
Efficient Policy Space Response Oracles
Ming Zhou, Jingxiao Chen, Ying Wen, Weinan Zhang, Yaodong Yang, Yong Yu, and Jun Wang. 2022 · 2022
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Large Language Models Can Self-Improve. In The 2023 Conference on Empirical Methods in Natural Language Processing
Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2023 · 2023
Later among the works it cites.
TiZero: Mastering Multi-Agent Football with Curriculum Learning and Self-Play. In Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems . 67–76
Fanqi Lin, Shiyu Huang, Tim Pearce, Wenze Chen, and Wei-Wei Tu. 2023 · 2023
Later among the works it cites.
Team-PSRO for Learning Approximate TMECor in Large Team Games via Cooperative Reinforcement Learning. In Thirty-seventh Conference on Neural Information Processing Systems
Stephen Marcus McAleer, Gabriele Farina, Gaoyue Zhou, Mingzhi Wang, Yaodong Yang, and Tuomas Sandholm. 2023 · 2023
Later among the works it cites.
Self-play reinforcement learning guides protein engineering
Yi Wang, Hui Tang, Lichao Huang, Lulu Pan, Lixiang Yang, Huanming Yang, Feng Mu, and Meng Yang. 2023 · 2023
Later among the works it cites.
Fictitious Cross-Play: Learning Global Nash Equilibrium in Mixed Cooperative-Competitive Games. In Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems . 1053–1061
Zelai Xu, Yancheng Liang, Chao Yu, Yu Wang, and Yi Wu. 2023 · 2023
Later among the works it cites.
Policy space diversity for non-transitive games
Jian Yao, Weiming Liu, Haobo Fu, Yaodong Yang, Stephen McAleer, Qiang Fu, and Wei Yang. 2023 · 2023
Later among the works it cites.
Policy space response oracles: a survey. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence . 7951–7961
Ariyan Bighashdel, Yongzhao Wang, Stephen McAleer, Rahul Savani, and Frans A Oliehoek. 2024 · 2024
Closest in time.
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models. In International Conference on Machine Learning . PMLR, 6621–6642
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu. 2024 · 2024
Closest in time.
Self-playing Adversarial Language Game Enhances LLM Reasoning. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
Pengyu Cheng, Tianhao Hu, Han Xu, Zhisong Zhang, Yong Dai, Lei Han, nan du, and Xiaolong Li. 2024 · 2024
Closest in time.
Human-compatible driving agents through data-regularized self-play reinforcement learning. In Reinforcement Learning Conference
Daphne Cornelisse and Eugene Vinitsky. 2024 · 2024
Closest in time.
Learning agile soccer skills for a bipedal robot with deep reinforcement learning
Tuomas Haarnoja, Ben Moran, Guy Lever, Sandy H Huang, Dhruva Tirumala, Jan Humplik, Markus Wulfmeier, Saran Tunyasuvunakool, Noah Y Siegel, Roland Hafner, et al · 2024
Closest in time.
A robust and opponent-aware league training method for starcraft ii
Ruozi Huang, Xipeng Wu, Hongsheng Yu, Zhong Fan, Haobo Fu, Qiang Fu, and Wei Yang. 2024 · 2024
Closest in time.
A survey on algorithms for Nash equilibria in finite normal-form games
Hanyu Li, Wenhan Huang, Zhijian Duan, David Henry Mguni, Kun Shao, Jun Wang, and Xiaotie Deng. 2024 · 2024
Closest in time.
Nash Learning from Human Feedback. In International Conference on Machine Learning . PMLR, 36743–36768
Remi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Zhaohan Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Côme Fiegel, et al · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback. In Proceedings of the 41st International Conference on Machine Learning . 47345–47377
Gokul Swamy, Christoph Dann, Rahul Kidambi, Zhiwei Steven Wu, and Alekh Agarwal. 2024 · 2024
Closest in time.
ASP: Learn a Universal Neural Solver!
Chenguang Wang, Zhouliang Yu, Stephen McAleer, Tianshu Yu, and Yaodong Yang. 2024 · 2024
Closest in time.
Self-Play Preference Optimization for Language Model Alignment. In ICML 2024 Workshop on Theoretical Foundations of Foundation Models
Yue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji, Yiming Yang, and Quanquan Gu. 2024 · 2024
Closest in time.
Mqe: Unleashing the power of interaction with multi-agent quadruped environment. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 5918–5924
Ziyan Xiong, Bo Chen, Shiyu Huang, Wei-Wei Tu, Zhaofeng He, and Yang Gao. 2024 · 2024
Closest in time.
Language agents with reinforcement learning for strategic play in the Werewolf game. In Proceedings of the 41st International Conference on Machine Learning . 55434–55464
Zelai Xu, Chao Yu, Fei Fang, Yu Wang, and Yi Wu. 2024 · 2024
Closest in time.
Self-rewarding language models
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. 2024 · 2024
Closest in time.
Achieving human level competitive robot table tennis. In 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 74–82
David B DAmbrosio, Saminda Abeyruwan, Laura Graesser, Atil Iscen, Heni Ben Amor, Alex Bewley, Barney J Reed, Krista Reymann, Leila Takayama, Yuval Tassa, et al · 2025
Closest in time.
Toward Real-World Cooperative and Competitive Soccer with Quadrupedal Robot Teams
Zhi Su, Yuman Gao, Emily Lukas, Yunfei Li, Jiaze Cai, Faris Tulbah, Fei Gao, Chao Yu, Zhongyu Li, Yi Wu, et al · 2025
Closest in time.
Hitter: A humanoid table tennis robot via hierarchical planning and learning
Zhi Su, Bike Zhang, Nima Rahmanian, Yuman Gao, Qiayuan Liao, Caitlin Regan, Koushil Sreenath, and S Shankar Sastry. 2025b · 2025
Closest in time.
Empirical Game Theoretic Analysis: A Survey
Michael P Wellman, Karl Tuyls, and Amy Greenwald. 2025 · 2025
Closest in time.
Zelai Xu, Wanjun Gu, Chao Yu, Yi Wu, and Yu Wang. 2025a · 2025
Closest in time.
Volleybots: A testbed for multi-drone volleyball game combining motion control and strategic play
Zelai Xu, Ruize Zhang, Chao Yu, Huining Yuan, Xiangmin Yi, Shilong Ji, Chuqi Wang, Wenhao Tang, Feng Gao, Wenbo Ding, et al · 2025
Closest in time.
Mastering Multi-Drone Volleyball through Hierarchical Co-Self-Play Reinforcement Learning
Ruize Zhang, Sirui Xiang, Zelai Xu, Feng Gao, Shilong Ji, Wenhao Tang, Wenbo Ding, Chao Yu, and Yu Wang. 2025 · 2025
Closest in time.