Fetching the paper…
Reading the bibliography…
Counterfactual Regret Minimization (CFR) is the most successful algorithm for finding approximate Nash equilibria in imperfect information games.
Random sampling with a reservoir
Vitter, J. S · 1985
Earlier work this paper cites.
Bayes’ bluff: Opponent modelling in poker
Southey, F., Bowling, M. P., Larson, B., Piccione, C., Burch, N., Billings, D., and Rayner, C · 2005
Earlier work this paper cites.
Regret minimization in games with incomplete information
Zinkevich, M., Johanson, M., Bowling, M., and Piccione, C · 2008
Earlier work this paper cites.
Monte carlo sampling for regret minimization in extensive games
Lanctot, M., Waugh, K., Zinkevich, M., and Bowling, M · 2009
Earlier work this paper cites.
Efficient monte carlo counterfactual regret minimization in games with many player actions
Burch, N., Lanctot, M., Szafron, D., and Gibson, R. G · 2012
Earlier work this paper cites.
Finding optimal abstract strategies in extensive-form games
Johanson, M., Bard, N., Burch, N., and Bowling, M · 2012
Earlier work this paper cites.
Potential-aware imperfect-recall abstraction with earth mover’s distance in imperfect-information games
Ganzfried, S. and Sandholm, T · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Solving large imperfect information games using cfr+
Tammelin, O · 2014
Earlier work this paper cites.
Heads-up limit hold’em poker is solved
Bowling, M., Burch, N., Johanson, M., and Tammelin, O · 2015
Cited alongside, same era.
Hierarchical abstraction, distributed equilibrium computation, and post-processing, with application to a champion no-limit texas hold’em agent
Brown, N., Ganzfried, S., and Sandholm, T · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Solving heads-up limit texas hold’em
Tammelin, O., Burch, N., Johanson, M., and Bowling, M · 2015
Cited alongside, same era.
Solving games with functional regret estimation
Waugh, K., Morrill, D., Bagnell, J. A., and Bowling, M · 2015
Cited alongside, same era.
Deep reinforcement learning from self-play in imperfect-information games
Regret minimization for partially observable deep reinforcement learning
Jin, P. H., Levine, S., and Keutzer, K · 2017
Later among the works it cites.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Pérolat, J., Silver, D., and Graepel, T · 2017
Later among the works it cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Moravčík, M., Schmid, M., Burch, N., Lisỳ, V., Morrill, D., Bard, N., Davis, T., Waugh, K., Johanson, M., and Bowling, M · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Heinrich, J. and Silver, D · 2016
Cited alongside, same era.
Equilibrium approximation quality of current no-limit poker bots
Lisy, V. and Bowling, M · 2016
Cited alongside, same era.
Learning values across many orders of magnitude
van Hasselt, H. P., Guez, A., Hessel, M., Mnih, V., and Silver, D · 2016
Cited alongside, same era.
Safe and nested subgame solving for imperfect-information games
Brown, N. and Sandholm, T · 2017
Cited alongside, same era.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T
Cited in the paper.
Solving imperfect-information games via discounted regret minimization
Brown, N. and Sandholm, T
Cited in the paper.
Deep counterfactual regret minimization
Brown, N., Lerer, A., Gross, S., and Sandholm, T
Cited in the paper.
Later among the works it cites.
Double neural counterfactual regret minimization
Hui, L., Kailiang, H., Zhibang, G., Tao, J., Yuan, Q., and Le, S · 2018
Later among the works it cites.
Schmid, M., Burch, N., Lanctot, M., Moravcik, M., Kadlec, R., and Bowling, M · 2018
Later among the works it cites.
Actor-critic policy optimization in partially observable multiagent environments
Srinivasan, S., Lanctot, M., Zambaldi, V., Pérolat, J., Tuyls, K., Munos, R., and Bowling, M · 2018
Later among the works it cites.