Fetching the paper…
Reading the bibliography…
Extensive-form games (EFGs) are a common model of multi-agent interactions with imperfect information.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
A Course in Game Theory
Martin J. Osborne and Ariel Rubinstein · 1994
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
Bayes’ bluff: Opponent modelling in poker
Finnegan Southey, Michael Bowling, Bryce Larson, Carmelo Piccione, Neil Burch, Darse Billings, and Chris Rayner · 2005
Earlier work this paper cites.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione · 2008
Earlier work this paper cites.
Natural actor–critic algorithms
Shalabh Bhatnagar, Richard S Sutton, Mohammad Ghavamzadeh, and Mark Lee · 2009
Earlier work this paper cites.
Monte Carlo sampling for regret minimization in extensive games
Marc Lanctot, Kevin Waugh, Martin Zinkevich, and Michael Bowling · 2009
Earlier work this paper cites.
Accelerating best response calculation in large extensive games
Michael Johanson, Kevin Waugh, Michael Bowling, , and Martin Zinkevich · 2011
Earlier work this paper cites.
Measuring the size of large no-limit poker games
Michael Johanson · 2013
Earlier work this paper cites.
Solving imperfect information games using decomposition
Neil Burch, Michael Johanson, and Michael Bowling · 2014
Earlier work this paper cites.
Heads-up limit hold’em poker is solved
Michael Bowling, Neil Burch, Michael Johanson, and Oskari Tammelin · 2015
Cited alongside, same era.
Faster first-order methods for extensive-form game solving
Christian Kroer, Kevin Waugh, Fatma Kilinç-Karzan, and Tuomas Sandholm · 2015
Cited alongside, same era.
Solving heads-up limit Texas hold’em
Oskari Tammelin, Neil Burch, Michael Johanson, and Michael Bowling · 2015
Cited alongside, same era.
Timeability of extensive-form games
Sune K. Jakobsen, Troels B. Sørensen, and Vincent Conitzer · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Cited alongside, same era.
Depth-limited solving for imperfect-information games
Noam Brown, Tuomas Sandholm, and Brandon Amos · 2018
Later among the works it cites.
AIVAT: A new variance reduction technique for agent evaluation in imperfect information games
Neil Burch, Martin Schmid, Matej Moravcik, Dustin Morrill, and Michael Bowling · 2018
Later among the works it cites.
Action-dependent control variates for policy optimization via Stein identity
Hao Liu, Yihao Feng, Yi Mao, Dengyong Zhou, Jian Peng, and Qiang Liu · 2018
Later among the works it cites.
Actor-critic policy optimization in partially observable multiagent environments
Sriram Srinivasan, Marc Lanctot, Vinicius Zambaldi, Julien Pérolat, Karl Tuyls, Rémi Munos, and Michael Bowling · 2018
Later among the works it cites.
The mirage of action-dependent baselines in reinforcement learning
George Tucker, Surya Bhupatiraju, Shixiang Gu, Richard Turner, Zoubin Ghahramani, and Sergey Levine · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data-efficient off-policy policy evaluation for reinforcement learning
Philip S. Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Safe and nested subgame solving for imperfect-information games
Noam Brown and Tuomas Sandholm · 2017
Cited alongside, same era.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisý, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael H. Bowling · 2017
Cited alongside, same era.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2018
Cited alongside, same era.
Efficient Monte Carlo counterfactual regret minimization in games with many player actions
Richard Gibson, Neil Burch, Marc Lanctot, and Duane Szafron
Cited in the paper.
Generalized sampling and variance in counterfactual regret minimization
Richard Gibson, Marc Lanctot, Neil Burch, Duane Szafron, and Michael Bowling
Cited in the paper.
Variance reduction for policy gradient with action-dependent factorized baselines
Cathy Wu, Aravind Rajeswaran, Yan Duan, Vikash Kumar, Alexandre M. Bayen, Sham M. Kakade, Igor Mordatch, and Pieter Abbeel · 2018
Later among the works it cites.
Lazy-CFR: a fast regret minimization algorithm for extensive games with imperfect information
Yichi Zhou, Tongzheng Ren, Jialian Li, Dong Yan, and Jun Zhu · 2018
Later among the works it cites.
CFR+, 2014
Computer Poker Research Group, University of Alberta and Oskari Tammelin · 2019
Closest in time.
Variance reduction in monte carlo counterfactual regret minimization (VR-MCCFR) for extensive form games using baselines
Martin Schmid, Neil Burch, Marc Lanctot, Matej Moravcik, Rudolf Kadlec, and Michael Bowling · 2019
Closest in time.
Monte Carlo continual resolving for online strategy computation in imperfect information games
Michal Šustr, Vojtěch Kovařík, and Viliam Lisý · 2019
Closest in time.