Fetching the paper…
Reading the bibliography…
Learning strategies for imperfect information games from samples of interaction is a challenging problem.
Options: A monte carlo approach
Phelim P Boyle · 1977
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R.J. Williams · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
S. Hart and A. Mas-Colell · 2000
Earlier work this paper cites.
Bayes’ bluff: Opponent modelling in poker
Finnegan Southey, Michael H. Bowling, Bryce Larson, Carmelo Piccione, Neil Burch, Darse Billings, and D. Chris Rayner · 2005
Earlier work this paper cites.
Regret minimization in games with incomplete information
M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione · 2008
Earlier work this paper cites.
Monte Carlo sampling for regret minimization in extensive games
Marc Lanctot, Kevin Waugh, Martin Zinkevich, and Michael Bowling · 2009
Earlier work this paper cites.
Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations
Y. Shoham and K. Leyton-Brown · 2009
Earlier work this paper cites.
Accelerating best response calculation in large extensive games
Michael Johanson, Michael Bowling, Kevin Waugh, and Martin Zinkevich · 2011
Earlier work this paper cites.
Variance reduction in Monte-Carlo tree search
Joel Veness, Marc Lanctot, and Michael Bowling · 2011
Earlier work this paper cites.
Generalized sampling and variance in counterfactual regret minimization
Richard Gibson, Marc Lanctot, Neil Burch, Duane Szafron, and Michael Bowling · 2012
Earlier work this paper cites.
Efficient nash equilibrium approximation through Monte Carlo counterfactual regret minimization
Michael Johanson, Nolan Bard, Marc Lanctot, Richard Gibson, and Michael Bowling · 2012
Earlier work this paper cites.
Monte Carlo Sampling and Regret Minimization for Equilibrium Computation and Decision-Making in Large Extensive Form Games
Marc Lanctot · 2013
Cited alongside, same era.
Monte Carlo theory, methods and examples
Art B. Owen · 2013
Cited alongside, same era.
Solving imperfect information games using decomposition
Neil Burch, Michael Johanson, and Michael Bowling · 2014
Cited alongside, same era.
Quality-based rewards for Monte-Carlo tree search simulations
Tom Pepels, Mandy J.W. Tak, Marc Lanctot, and Mark H.M. Winands · 2014
Cited alongside, same era.
Heads-up Limit Hold’em Poker is solved
Michael Bowling, Neil Burch, Michael Johanson, and Oskari Tammelin · 2015
Cited alongside, same era.
Fictitious self-play in extensive-form games
Johannes Heinrich, Marc Lanctot, and David Silver · 2015
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Perolat, David Silver, and Thore Graepel · 2017
Later among the works it cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisý, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Reinforcement Learning: An Introduction
R. Sutton and A. Barto · 2017
Later among the works it cites.
Rudder: Return decomposition for delayed rewards
Jose A. Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Unterthiner, and Sepp Hochreiter · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Solving heads-up limit texas hold’em
Oskari Tammelin, Neil Burch, Michael Johanson, and Michael Bowling · 2015
Cited alongside, same era.
Solving games with functional regret estimation
Kevin Waugh, Dustin Morrill, J. Andrew Bagnell, and Michael Bowling · 2015
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2017
Cited alongside, same era.
Time and Space: Why Imperfect Information Games are Hard
Neil Burch · 2017
Cited alongside, same era.
Learning with opponent-learning awareness
Jakob N. Foerster, Richard Y. Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2018
Closest in time.
Aivat: A new variance reduction technique for agent evaluation in imperfect information games, 2018
Neil Burch, Martin Schmid, Matej Moravcik, Dustin Morill, and Michael Bowling · 2018
Closest in time.
Kuhn poker — Wikipedia, the free encyclopedia, 2018
Kuhn poker · 2018
Closest in time.
Action-dependent control variates for policy optimization via stein identity
Hao Liu, Yihao Feng, Yi Mao, Dengyong Zhou, Jian Peng, and Qiang Liu · 2018
Closest in time.
The mirage of action-dependent baselines in reinforcement learning
George Tucker, Surya Bhupatiraju, Shixiang Gu, Richard E Turner, Zoubin Ghahramani, and Sergey Levine · 2018
Closest in time.
Variance reduction for policy gradient with action-dependent factorized baselines
Cathy Wu, Aravind Rajeswaran, Yan Duan, Vikash Kumar, Alexandre M. Bayen, Sham Kakade, Igor Mordatch, and Pieter Abbeel · 2018
Closest in time.