Fetching the paper…
Reading the bibliography…
As a challenging multi-player card game, DouDizhu has recently drawn much attention for analyzing competition and collaboration in imperfect-information games.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione · 2007
Earlier work this paper cites.
Automating collusion detection in sequential games
Parisa Mazrooei, Christopher Archibald, and Michael Bowling · 2013
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Earlier work this paper cites.
Solving games with functional regret estimation
Kevin Waugh, Dustin Morrill, James Andrew Bagnell, and Michael Bowling · 2015
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Jakob N. F., Yannis M. A., Nando de F., and Shimon W · 2016
Earlier work this paper cites.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan L., Yi W., Aviv T., Jean H., Pieter A., and Igor M · 2017
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolat, David Silver, and Thore Graepel · 2017
Earlier work this paper cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisỳ, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Cited alongside, same era.
Horovod: fast and easy distributed deep learning in tensorflow
Alexander Sergeev and Mike Del Balso · 2018
Cited alongside, same era.
Actor-critic policy optimization in partially observable multiagent environments
Sriram Srinivasan, Marc Lanctot, Vinicius Zambaldi, Julien Pérolat, Karl Tuyls, Rémi Munos, and Michael Bowling · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Improving policies via search in cooperative partially observable games
Adam Lerer, Hengyuan Hu, Jakob Foerster, and Noam Brown · 2020
Later among the works it cites.
Suphx: Mastering mahjong with deep reinforcement learning
Junjie Li, Sotetsu Koyamada, Qiwei Ye, Guoqing Liu, Chao Wang, Ruihan Yang, Li Zhao, Tao Qin, Tie-Yan Liu, and Hsiao-Wuen Hon · 2020
Later among the works it cites.
Dream: Deep regret minimization with advantage baselines and model-free learning
Eric Steinberger, Adam Lerer, and Noam Brown · 2020
Later among the works it cites.
Mastering complex control in moba games with deep reinforcement learning
Deheng Ye, Zhao Liu, Mingfei Sun, Bei Shi, Peilin Zhao, Hao Wu, Hongsheng Yu, Shaojie Yang, Xipeng Wu, Qingwei Guo, et al · 2020
Later among the works it cites.
Universal trading for order execution with oracle policy distillation
Yuchen Fang, Kan Ren, Weiqing Liu, Dong Zhou, Weinan Zhang, Jiang Bian, Yong Yu, and Tie-Yan Liu · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemyslaw Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Cited alongside, same era.
Deep counterfactual regret minimization
Noam Brown, Adam Lerer, Sam Gross, and Tuomas Sandholm · 2019
Cited alongside, same era.
Deltadou: Expert-level doudizhu ai through self-play
Qiqi Jiang, Kuangzheng Li, Boyao Du, Hao Chen, and Hai Fang · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Combinational q-learning for dou di zhu
Yang You, Liangwei Li, Baisong Guo, Weiming Wang, and Cewu Lu · 2019
Cited alongside, same era.
Rlcard: A toolkit for reinforcement learning in card games
Daochen Zha, Kwei-Herng Lai, Yuanpu Cao, Songyi Huang, Ruzhe Wei, Junyu Guo, and Xia Hu · 2019
Cited alongside, same era.
The advantage regret-matching actor-critic
Audrūnas Gruslys, Marc Lanctot, Rémi Munos, Finbarr Timbers, Martin Schmid, Julien Perolat, Dustin Morrill, Vinicius Zambaldi, Jean-Baptiste Lespiau, John Schultz, et al · 2020
Cited alongside, same era.
Later among the works it cites.
Actor-critic policy optimization in a large-scale imperfect-information game
Haobo Fu, Weiming Liu, Shuang Wu, Yijia Wang, Tao Yang, Kai Li, Junliang Xing, Bin Li, Bo Ma, Qiang Fu, et al · 2021
Later among the works it cites.
A survey of nash equilibrium strategy solving based on cfr
Huale Li, Xuan Wang, Fengwei Jia, Yifan Li, and Qian Chen · 2021
Later among the works it cites.
Douzero: Mastering doudizhu with self-play deep reinforcement learning
Daochen Zha, Jingru Xie, Wenye Ma, Sheng Zhang, Xiangru Lian, Xia Hu, and Ji Liu · 2021
Later among the works it cites.
Combining tree search and action prediction for state-of-the-art performance in doudizhu
Yunsheng Zhang, Dong Yan, Bei Shi, Haobo Fu, Qiang Fu, Hang Su, Jun Zhu, and Ning Chen · 2021
Later among the works it cites.
Alphaholdem: High-performance artificial intelligence for heads-up no-limit texas hold’em from end-to-end reinforcement learning
Enmin Zhao, Renye Yan, Jinqiu Li, Kai Li, and Junliang Xing · 2022
Closest in time.
Douzero+: Improving doudizhu ai by opponent modeling and coach-guided learning
Youpeng Zhao, Jian Zhao, Xunhan Hu, Wengang Zhou, and Houqiang Li · 2022
Closest in time.