Fetching the paper…
Reading the bibliography…
Counterfactual Regret Minimization (CFR) is the leading framework for solving large imperfect-information games.
Equilibrium points in n-person games
Nash, J · 1950
Earlier work this paper cites.
Iterative solutions of games by fictitious play
Brown, G. W · 1951
Earlier work this paper cites.
Random sampling with a reservoir
Vitter, J. S · 1985
Earlier work this paper cites.
The weighted majority algorithm
Littlestone, N. and Warmuth, M. K · 1994
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Hart, S. and Mas-Colell, A · 2000
Earlier work this paper cites.
Clustering with bregman divergences
Banerjee, A., Merugu, S., Dhillon, I. S., and Ghosh, J · 2005
Earlier work this paper cites.
Prediction, learning, and games
Cesa-Bianchi, N. and Lugosi, G · 2006
Earlier work this paper cites.
Regret minimization in games with incomplete information
Zinkevich, M., Johanson, M., Bowling, M. H., and Piccione, C · 2007
Earlier work this paper cites.
A parameter-free hedging algorithm
Chaudhuri, K., Freund, Y., and Hsu, D. J · 2009
Earlier work this paper cites.
Monte Carlo sampling for regret minimization in extensive games
Lanctot, M., Waugh, K., Zinkevich, M., and Bowling, M · 2009
Earlier work this paper cites.
Abstraction in large extensive games
Waugh, K · 2009
Earlier work this paper cites.
Smoothing techniques for computing Nash equilibria of sequential games
Hoda, S., Gilpin, A., Peña, J., and Sandholm, T · 2010
Earlier work this paper cites.
Generalized sampling and variance in counterfactual regret minimization
Gibson, R., Lanctot, M., Burch, N., Szafron, D., and Bowling, M · 2012
Earlier work this paper cites.
Efficient nash equilibrium approximation through monte carlo counterfactual regret minimization
Johanson, M., Bard, N., Lanctot, M., Gibson, R., and Bowling, M · 2012
Earlier work this paper cites.
Evaluating state-space abstractions in extensive-form games
Johanson, M., Burch, N., Valenzano, R., and Bowling, M · 2013
Earlier work this paper cites.
Monte carlo sampling and regret minimization for equilibrium computation and decision-making in large extensive form games
Lanctot, M · 2013
Earlier work this paper cites.
Potential-aware imperfect-recall abstraction with earth mover’s distance in imperfect-information games
Ganzfried, S. and Sandholm, T · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Heads-up limit hold’em poker is solved
Bowling, M., Burch, N., Johanson, M., and Tammelin, O · 2015
Cited alongside, same era.
Hierarchical abstraction, distributed equilibrium computation, and post-processing, with application to a champion no-limit texas hold’em agent
Brown, N., Ganzfried, S., and Sandholm, T · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Regret minimization for partially observable deep reinforcement learning
Jin, P. H., Levine, S., and Keutzer, K · 2017
Later among the works it cites.
Playing FPS games with deep reinforcement learning
Lample, G. and Chaplot, D. S · 2017
Later among the works it cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Moravčík, M., Schmid, M., Burch, N., Lisý, V., Morrill, D., Bard, N., Davis, T., Waugh, K., Johanson, M., and Bowling, M · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Solving heads-up limit texas hold’em
Tammelin, O., Burch, N., Johanson, M., and Bowling, M · 2015
Cited alongside, same era.
Solving games with functional regret estimation
Waugh, K., Morrill, D., Bagnell, D., and Bowling, M · 2015
Cited alongside, same era.
Deep reinforcement learning from self-play in imperfect-information games
Heinrich, J. and Silver, D · 2016
Cited alongside, same era.
Compact CFR
Jackson, E. G · 2016
Cited alongside, same era.
Robust Strategies and Counter-Strategies: From Superhuman to Optimal Play
Johanson, M. B · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Using Regret Estimation to Solve Games Compactly
Morrill, D. R · 2016
Cited alongside, same era.
Depth-limited solving for imperfect-information games
Brown, N., Sandholm, T., and Amos, B · 2018
Closest in time.
Revisiting cfr+ and alternating updates
Burch, N., Moravcik, M., and Schmid, M · 2018
Closest in time.
Online convex optimization for sequential decision processes and extensive-form games
Farina, G., Kroer, C., and Sandholm, T · 2018
Closest in time.
Double neural counterfactual regret minimization
Li, H., Hu, K., Ge, Z., Jiang, T., Qi, Y., and Song, L · 2018
Closest in time.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Closest in time.
Actor-critic policy optimization in partially observable multiagent environments
Srinivasan, S., Lanctot, M., Zambaldi, V., Pérolat, J., Tuyls, K., Munos, R., and Bowling, M · 2018
Closest in time.
Solving imperfect-information games via discounted regret minimization
Brown, N. and Sandholm, T · 2019
Closest in time.
Stable-predictive optimistic counterfactual regret minimization
Farina, G., Kroer, C., Brown, N., and Sandholm, T · 2019
Closest in time.
Variance reduction in monte carlo counterfactual regret minimization (VR-MCCFR) for extensive form games using baselines
Schmid, M., Burch, N., Lanctot, M., Moravcik, M., Kadlec, R., and Bowling, M · 2019
Closest in time.
Single deep counterfactual regret minimization
Steinberger, E · 2019
Closest in time.