Fetching the paper…
Reading the bibliography…
From the very dawn of the field, search with value functions was a fundamental concept of computer games research.
Single deep counterfactual regret minimization
Steinberger, E. (2019) · 1901
Earlier work this paper cites.
Rethinking formal models of partially observable multiagent decision making
Kovařík, V., Schmid, M., Burch, N., Bowling, M., and Lisỳ, V. (2019) · 1906
Earlier work this paper cites.
Value functions for depth-limited solving in imperfect-information games
Kovarık, V., Seitz, D., Lisỳ, V., Rudolf, J., Sun, S., and Ha, K. (2020) · 1906
Earlier work this paper cites.
Problems with the efg formalism: a solution attempt using observations
Kovařík, V. and Lisý, V. (2019) · 1906
Earlier work this paper cites.
Value functions for depth-limited solving in imperfect-information games beyond poker
Seitz, D., Kovařík, V., Lisỳ, V., Rudolf, J., Sun, S., and Ha, K. (2019) · 1906
Earlier work this paper cites.
A generalized training approach for multiagent learning
Muller, P., Omidshafiei, S., Rowland, M., Tuyls, K., Perolat, J., Liu, S., Hennes, D., Marris, L., Lanctot, M., Hughes, E., et al. (2019) · 1909
Earlier work this paper cites.
Optimistic regret minimization for extensive-form games via dilated distance-generating functions
Farina, G., Kroer, C., and Sandholm, T. (2019b) · 1910
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al. (2019) · 1912
Earlier work this paper cites.
Online convex optimization for sequential decision processes and extensive-form games
Farina, G., Kroer, C., and Sandholm, T. (2019a) · 1925
Earlier work this paper cites.
Zur theorie der gesellschaftsspiele
Neumann, J. v. (1928) · 1928
Earlier work this paper cites.
A simplified two-person poker
Kuhn, H. W. (1950) · 1950
Earlier work this paper cites.
Equilibrium points in n-person games
Nash, J. F. et al. (1950) · 1950
Earlier work this paper cites.
Xxii. programming a computer for playing chess
Shannon, C. E. (1950) · 1950
Earlier work this paper cites.
Iterative solution of games by fictitious play
Brown, G. W. (1951) · 1951
Earlier work this paper cites.
Linear programming and the theory of games
Gale, D., Kuhn, H., and Tucker, A. (1951) · 1951
Earlier work this paper cites.
Non-cooperative games
Nash, J. (1951) · 1951
Earlier work this paper cites.
An iterative method of solving a game
Robinson, J. (1951) · 1951
Earlier work this paper cites.
Extensive games and the problem of information. kuhn hw, tucker aw, eds., contributions to the theory of games, vol ii, 193–216
Kuhn, H. (1953) · 1953
Earlier work this paper cites.
Theory of games and economic behavior
Morgenstern, O. and Von Neumann, J. (1953) · 1953
Earlier work this paper cites.
Stochastic games
Shapley, L. S. (1953) · 1953
Earlier work this paper cites.
Communication on the borel notes
Von Neumann, J. and Fréchet, M. (1953) · 1953
Earlier work this paper cites.
An analog of the minimax theorem for vector payoffs
Blackwell, D. et al. (1956) · 1956
Earlier work this paper cites.
A markovian decision process
Bellman, R. (1957) · 1957
Earlier work this paper cites.
Some studies in machine learning using the game of checkers
Samuel, A. L. (1959) · 1959
Earlier work this paper cites.
Mixed and behavior strategies in infinite extensive games
Aumann, R. J. (1961) · 1961
Earlier work this paper cites.
An example of human chess play in the light of chess playing programs
Newell, A. and Simon, H. A. (1964) · 1964
Earlier work this paper cites.
Optimal control of markov processes with incomplete state information
Astrom, K. J. (1965) · 1965
Earlier work this paper cites.
Prisoner’s dilemma: A study in conflict and cooperation
Rapoport, A., Chammah, A. M., and Orwant, C. J. (1965) · 1965
Earlier work this paper cites.
The game of chicken
Rapoport, A. and Chammah, A. M. (1966) · 1966
Earlier work this paper cites.
An analysis of alpha-beta pruning
Knuth, D. E. and Moore, R. W. (1975) · 1975
Earlier work this paper cites.
Learning a value analysis tool for agent evaluation
White, M. and Bowling, M. H. (2009) · 1981
Earlier work this paper cites.
An investigation of the causes of pathology in games
Nau, D. S. (1982) · 1982
Earlier work this paper cites.
Chess as the drosophila of ai
McCarthy, J. (1990) · 1990
Earlier work this paper cites.
Robust estimation of a location parameter
Huber, P. J. (1992) · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
The weighted majority algorithm
Littlestone, N. and Warmuth, M. K. (1994) · 1994
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L. (1994) · 1994
Earlier work this paper cites.
A course in game theory
Osborne, M. J. and Rubinstein, A. (1994) · 1994
Earlier work this paper cites.
On the complexity of the parity argument and other inefficient proofs of existence
Papadimitriou, C. H. (1994) · 1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
Tesauro, G. (1995) · 1995
Earlier work this paper cites.
The absent minded driver’s paradox: Synthesis and responses
Piccione, M., Rubinstein, A., et al. (1996) · 1996
Earlier work this paper cites.
The lack of a priori distinctions between learning algorithms
Wolpert, D. H. (1996) · 1996
Earlier work this paper cites.
Introduction to linear optimization
Bertsimas, D. and Tsitsiklis, J. N. (1997) · 1997
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Freund, Y. and Schapire, R. E. (1997) · 1997
Earlier work this paper cites.
On the interpretation of decision problems with imperfect recall
Piccione, M. and Rubinstein, A. (1997) · 1997
Earlier work this paper cites.
Finding optimal strategies for imperfect information games
Frank, I., Basin, D. A., and Matsubara, H. (1998) · 1998
Earlier work this paper cites.
On the rate of convergence of continuous-time fictitious play
Harris, C. (1998) · 1998
Earlier work this paper cites.
Tracking the best expert
Herbster, M. and Warmuth, M. K. (1998) · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R. (1998) · 1998
Earlier work this paper cites.
Introduction to reinforcement learning
Sutton, R. S., Barto, A. G., et al. (1998) · 1998
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Hart, S. and Mas-Colell, A. (2000) · 2000
Earlier work this paper cites.
A weakened form of fictitious play in two-person zero-sum games
Van der Genugten, B. (2000) · 2000
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
Bernstein, D. S., Givan, R., Immerman, N., and Zilberstein, S. (2002) · 2002
Cited alongside, same era.
Deep blue
Campbell, M., Hoane Jr, A. J., and Hsu, F.-h. (2002) · 2002
Cited alongside, same era.
Perolat, J., Munos, R., Lespiau, J.-B., Omidshafiei, S., Rowland, M., Ortega, P., Burch, N., Anthony, T., Balduzzi, D., De Vylder, B., et al. (2020) · 2002
Cited alongside, same era.
Multiagent learning in the presence of agents with limitations
Bowling, M. (2003) · 2003
Cited alongside, same era.
Planning in the presence of cost functions controlled by an adversary
McMahan, H. B., Gordon, G. J., and Blum, A. (2003) · 2003
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Later among the works it cites.
Search in imperfect information games using online monte carlo counterfactual regret minimization
Lanctot, M., Lisy, V., and Bowling, M. (2014) · 2014
Later among the works it cites.
Bounding the support size in extensive form games with imperfect information
Schmid, M., Moravcik, M., and Hladik, M. (2014) · 2014
Later among the works it cites.
Solving large imperfect information games using cfr+
Tammelin, O. (2014) · 2014
Later among the works it cites.
Monte Carlo tree search for games with hidden information and uncertainty
Whitehouse, D. (2014) · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zinkevich, M. (2003) · 2003
Cited alongside, same era.
Convex optimization
Boyd, S. and Vandenberghe, L. (2004) · 2004
Cited alongside, same era.
The essential turing
Copeland, B. J. (2004) · 2004
Cited alongside, same era.
Approximate exploitability: Learning a best response in large games
Timbers, F., Lockhart, E., Schmid, M., Lanctot, M., and Bowling, M. (2020) · 2004
Cited alongside, same era.
Bayes’ bluff: opponent modelling in poker
Southey, F., Bowling, M., Larson, B., Piccione, C., Burch, N., Billings, D., and Rayner, C. (2005) · 2005
Cited alongside, same era.
The mathematics of poker
Chen, B. and Ankenman, J. (2006) · 2006
Cited alongside, same era.
Settling the complexity of two-player nash equilibrium
Chen, X. and Deng, X. (2006) · 2006
Cited alongside, same era.
Heads-up limit hold’em poker is solved
Bowling, M., Burch, N., Johanson, M., and Tammelin, O. (2015) · 2015
Later among the works it cites.
Endgame solving in large imperfect-information games
Ganzfried, S. and Sandholm, T. (2015) · 2015
Later among the works it cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Later among the works it cites.
Fictitious self-play in extensive-form games
Heinrich, J., Lanctot, M., and Silver, D. (2015) · 2015
Later among the works it cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G. (2015) · 2015
Later among the works it cites.
Online monte carlo counterfactual regret minimization for search in imperfect information games
Lisý, V., Lanctot, M., and Bowling, M. (2015) · 2015
Later among the works it cites.
Automatic public state space abstraction in imperfect information games
Schmid, M., Moravcik, M., Hladik, M., and Gaukrodger, S. J. (2015) · 2015
Later among the works it cites.
Deep learning in neural networks: An overview
Schmidhuber, J. (2015) · 2015
Later among the works it cites.
Fast convergence of regularized learning in games
Syrgkanis, V., Agarwal, A., Luo, H., and Schapire, R. E. (2015) · 2015
Later among the works it cites.
Solving heads-up limit texas hold’em
Tammelin, O., Burch, N., Johanson, M., and Bowling, M. (2015) · 2015
Later among the works it cites.
Solving games with functional regret estimation
Waugh, K., Morrill, D., Bagnell, J. A., and Bowling, M. (2015) · 2015
Later among the works it cites.
Online Agent Modelling in Human-Scale Problems
Bard, N. D. (2016) · 2016
Later among the works it cites.
Reflections on the first man vs. machine no-limit texas hold’em competition
Ganzfried, S. (2016) · 2016
Later among the works it cites.
Timeability of extensive-form games
Jakobsen, S. K., Sørensen, T. B., and Conitzer, V. (2016) · 2016
Later among the works it cites.
Robust strategies and counter-strategies: from superhuman to optimal play
Johanson, M. B. (2016) · 2016
Later among the works it cites.
Refining subgames in large imperfect information games
Moravcik, M., Schmid, M., Ha, K., Hladik, M., and Gaukrodger, S. J. (2016) · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Later among the works it cites.
Safe and nested subgame solving for imperfect-information games
Brown, N. and Sandholm, T. (2017) · 2017
Later among the works it cites.
An algorithm for constructing and solving imperfect recall abstractions of large extensive-form games
Čermák, J., Bošansky, B., and Lisy, V. (2017) · 2017
Later among the works it cites.
Solving for best responses and equilibria in extensive-form games with reinforcement learning methods
Greenwald, A., Li, J., and Sodomka, E. (2017) · 2017
Later among the works it cites.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Pérolat, J., Silver, D., and Graepel, T. (2017) · 2017
Later among the works it cites.
Eqilibrium approximation quality of current no-limit poker bots
Lisy, V. and Bowling, M. (2017) · 2017
Later among the works it cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Moravčík, M., Schmid, M., Burch, N., Lisỳ, V., Morrill, D., Bard, N., Davis, T., Waugh, K., Johanson, M., and Bowling, M. (2017) · 2017
Later among the works it cites.
Rudder: Return decomposition for delayed rewards
Arjona-Medina, J. A., Gillhofer, M., Widrich, M., Unterthiner, T., Brandstetter, J., and Hochreiter, S. (2018) · 2018
Later among the works it cites.
Multiplicative weights update in zero-sum games
Bailey, J. P. and Piliouras, G. (2018) · 2018
Later among the works it cites.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T. (2018) · 2018
Later among the works it cites.
Time and space: Why imperfect information games are hard
Burch, N. (2018) · 2018
Later among the works it cites.
Aivat: A new variance reduction technique for agent evaluation in imperfect information games
Burch, N., Schmid, M., Moravcik, M., Morill, D., and Bowling, M. (2018) · 2018
Later among the works it cites.
Online convex optimization for sequential decision processes and extensive-form games
Farina, G., Kroer, C., and Sandholm, T. (2018) · 2018
Later among the works it cites.
Double neural counterfactual regret minimization
Li, H., Hu, K., Ge, Z., Jiang, T., Qi, Y., and Song, L. (2018) · 2018
Later among the works it cites.
Monte carlo continual resolving for online strategy computation in imperfect information games
Sustr, M., Kovarik, V., and Lisy, V. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
The mirage of action-dependent baselines in reinforcement learning
Tucker, G., Bhupatiraju, S., Gu, S., Turner, R. E., Ghahramani, Z., and Levine, S. (2018) · 2018
Later among the works it cites.
Revisiting cfr+ and alternating updates
Burch, N., Moravcik, M., and Schmid, M. (2019) · 2019
Later among the works it cites.
Variance reduction in monte carlo counterfactual regret minimization (vr-mccfr) for extensive form games using baselines
Schmid, M., Burch, N., Lanctot, M., Moravcik, M., Kadlec, R., and Bowling, M. (2019) · 2019
Later among the works it cites.
Finding friend and foe in multi-agent games
Serrino, J., Kleiman-Weiner, M., Parkes, D. C., and Tenenbaum, J. (2019) · 2019
Later among the works it cites.
Combining deep reinforcement learning and search for imperfect-information games
Brown, N., Bakhtin, A., Lerer, A., and Gong, Q. (2020) · 2020
Later among the works it cites.
Low-variance and zero-variance baselines for extensive-form games
Davis, T., Schmid, M., and Bowling, M. (2020) · 2020
Later among the works it cites.
Leela zero
Pascutto, G.-C. (2019) · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al. (2020) · 2020
Later among the works it cites.
Man vs. machine poker challenge
Hedberg, S. (2007) · 2021
Closest in time.
Team polk’s bryan pellegrino talks about his ai research and how it helped formulate strategies to win $1.2 million
Hedberg, S. (2021) · 2021
Closest in time.
The international federation of match poker
IFMP (2021) · 2021
Closest in time.
Safe search for stackelberg equilibria in extensive-form games
Ling, C. K. and Brown, N. (2021) · 2021
Closest in time.
Xdo: A double oracle algorithm for extensive-form games
McAleer, S., Lanier, J., Baldi, P., and Fox, R. (2021) · 2021
Closest in time.
Sound algorithms in imperfect information games
Šustr, M., Schmid, M., Moravčík, M., Burch, N., Lanctot, M., and Bowling, M. (2021) · 2021
Closest in time.