Fetching the paper…
Reading the bibliography…
Two-player, constant-sum games are well studied in the literature, but there has been limited progress outside of this setting.
Open-ended learning in symmetric zero-sum games
Balduzzi, D., Garnelo, M., Bachrach, Y., Czarnecki, W. M., Pérolat, J., Jaderberg, M., and Graepel, T · 1901
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., de Oliveira Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 1912
Earlier work this paper cites.
Uber die Abgrenzung der Eigenwerte einer Matrix
Gerschgorin, S · 1931
Earlier work this paper cites.
Contributions to the theory of statistical estimation and testing hypotheses
Wald, A · 1939
Earlier work this paper cites.
Statistical decision functions which minimize the maximum risk
Wald, A · 1945
Earlier work this paper cites.
A mathematical theory of communication
Shannon, C. E · 1948
Earlier work this paper cites.
A simplified two-person poker
Kuhn, H. W · 1950
Earlier work this paper cites.
Iterative solutions of games by fictitious play
Brown, G. W · 1951
Earlier work this paper cites.
Non-cooperative games
Nash, J · 1951
Earlier work this paper cites.
The foundations of statistics
Leonard J. Savage, J. W · 1954
Earlier work this paper cites.
Information theory and statistical mechanics
Jaynes, E. T · 1957
Earlier work this paper cites.
Quantification method of classification processes: Concept of structural a-entropy
Havrda, J., Charvat, F., and Havrda, J · 1967
Earlier work this paper cites.
The conjugate gradient method in extreme problem
Polyak, B · 1969
Earlier work this paper cites.
Subjectivity and correlation in randomized strategies
Aumann, R · 1974
Earlier work this paper cites.
Strategically zero-sum games: the class of games whose completely mixed equilibria cannot be improved upon
Moulin, H. and Vial, J.-P · 1978
Earlier work this paper cites.
A generalized conjugate gradient algorithm for solving a class of quadratic programming problems
O’Leary, D. P · 1980
Earlier work this paper cites.
A Rehabilitation of the Principle of Insufficient Reason
Sinn, H.-W · 1980
Earlier work this paper cites.
Classification and Regression Trees
Breiman, L., Friedman, J. H., Olshen, R. A., and Stone, C. J · 1984
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
A General Theory of Equilibrium Selection in Games , volume 1
Harsanyi, J. and Selten, R · 1988
Earlier work this paper cites.
Possible generalization of Boltzmann-Gibbs statistics
Tsallis, C · 1988
Earlier work this paper cites.
Game Theory
Fudenberg, D. and Tirole, J · 1991
Cited alongside, same era.
A limited memory algorithm for bound constrained optimization
Byrd, R., Lu, P., Nocedal, J., and Zhu, C · 1995
Cited alongside, same era.
Support-vector networks
Cortes, C. and Vapnik, V · 1995
Cited alongside, same era.
Microeconomic Theory
Mas-Colell, A., Whinston, M. D., and Green, J. R · 1995
Cited alongside, same era.
Planning in the presence of cost functions controlled by an adversary
McMahan, H. B., Gordon, G. J., and Blum, A · 2003
Cited alongside, same era.
On the geometry of Nash equilibria and correlated equilibria
Nau, R., Canovas, S. G., and Hansen, P · 2004
Cited alongside, same era.
Pattern Recognition and Machine Learning (Information Science and Statistics)
Unifying attribute splitting criteria of decision trees by Tsallis entropy
Wang, Y. and Xia, S · 2017
Later among the works it cites.
A rewriting system for convex optimization problems
Agrawal, A., Verschueren, R., Diamond, S., and Boyd, S · 2018
Later among the works it cites.
Re-evaluating evaluation
Balduzzi, D., Tuyls, K., Perolat, J., and Graepel, T · 2018
Later among the works it cites.
Superhuman AI for multiplayer poker
Brown, N. and Sandholm, T · 2019
Later among the works it cites.
Learning to correlate in multi-player general-sum sequential games, 2019
Celli, A., Marchesi, A., Bianchi, T., and Gatti, N · 2019
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Jaderberg, M., Czarnecki, W., Dunning, I., Marris, L., Lever, G., Castañeda, A., Beattie, C., Rabinowitz, N., Morcos, A., Ruderman, A., Sonnerat, N., Green, T., Deason, L., Leibo, J., Silver, D., Hassabis, D., Kavukcuoglu, K., and Graepel, T · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bishop, C. M · 2006
Cited alongside, same era.
Understanding and Using Linear Programming (Universitext)
Matouek, J. and Gärtner, B · 2006
Cited alongside, same era.
Maximum entropy correlated equilibria
Ortiz, L. E., Schapire, R. E., and Kakade, S. M · 2007
Cited alongside, same era.
Extensive-form correlated equilibrium: Definition and computational complexity
von Stengel, B. and Forges, F · 2008
Cited alongside, same era.
Robust Optimization
Ben-Tal, A., Ghaoui, L., and Nemirovski, A · 2009
Cited alongside, same era.
The complexity of computing a Nash equilibrium
Daskalakis, C., Goldberg, P., and Papadimitriou, C · 2009
Cited alongside, same era.
Later among the works it cites.
A brief review on different measures of entropy
Kaur, M. and Buttar, G · 2019
Later among the works it cites.
OpenSpiel: A framework for reinforcement learning in games
Lanctot, M., Lockhart, E., Lespiau, J.-B., Zambaldi, V., Upadhyay, S., Pérolat, J., Srinivasan, S., Timbers, F., Tuyls, K., Omidshafiei, S., Hennes, D., Morrill, D., Muller, P., Ewalds, T., Faulkner, R., Kramár, J., Vylder, B. D., Saeta, B., Bradbury, J., Ding, D., Borgeaud, S., Lai, M., Schrittwieser, J., Anthony, T., Hughes, E., Danihelka, I., and Ryan-Davis, J · 2019
Later among the works it cites.
α \alpha -rank: Multi-agent evaluation by evolution
Omidshafiei, S., Papadimitriou, C., Piliouras, G., Tuyls, K., Rowland, M., Lespiau, J.-B., Czarnecki, W. M., Lanctot, M., Perolat, J., and Munos, R · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W., Mathieu, M., Dudzik, A., Chung, J., Choi, D., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J., Jaderberg, M., and Silver, D · 2019
Later among the works it cites.
Learning to play no-press diplomacy with best response policy iteration, 2020
Anthony, T., Eccles, T., Tacchetti, A., Kramár, J., Gemp, I., Hudson, T. C., Porcel, N., Lanctot, M., Pérolat, J., Everett, R., Werpachowski, R., Singh, S., Graepel, T., and Bachrach, Y · 2020
Later among the works it cites.
No-regret learning dynamics for extensive-form correlated equilibrium, 2020
Celli, A., Marchesi, A., Farina, G., and Gatti, N · 2020
Later among the works it cites.
Human-level performance in no-press diplomacy via equilibrium search, 2020
Gray, J., Lerer, A., Bakhtin, A., and Brown, N · 2020
Later among the works it cites.
Human-agent cooperation in bridge bidding, 2020
Lockhart, E., Burch, N., Bard, N., Borgeaud, S., Eccles, T., Smaira, L., and Smith, R · 2020
Later among the works it cites.
Pipeline PSRO: A scalable approach for finding approximate Nash equilibria in large games
McAleer, S., Lanier, J., Fox, R., and Baldi, P · 2020
Later among the works it cites.
A generalized training approach for multiagent learning
Muller, P., Omidshafiei, S., Rowland, M., Tuyls, K., Perolat, J., Liu, S., Hennes, D., Marris, L., Lanctot, M., Hughes, E., Wang, Z., Lever, G., Heess, N., Graepel, T., and Munos, R · 2020
Later among the works it cites.
OSQP: an operator splitting solver for quadratic programs
Stellato, B., Banjac, G., Goulart, P., Bemporad, A., and Boyd, S · 2020
Later among the works it cites.
XDO: A double oracle algorithm for extensive-form games
McAleer, S., Lanier, J. B., Baldi, P., and Fox, R · 2021
Closest in time.
Hindsight and sequential rationality of correlated play
Morrill, D., D’Orazio, R., Sarfati, R., Lanctot, M., Wright, J. R., Greenwald, A., and Bowling, M · 2021
Closest in time.
Solving common-payoff games with approximate policy iteration, 2021
Sokota, S., Lockhart, E., Timbers, F., Davoodi, E., D’Orazio, R., Burch, N., Schmid, M., Bowling, M., and Lanctot, M · 2021
Closest in time.