Fetching the paper…
Reading the bibliography…
Tremendous advances have been made in multiagent reinforcement learning (MARL).
1901
Earlier work this paper cites.
1901
Earlier work this paper cites.
Jin, C., P. Netrapalli, and M. I. Jordan (2019b), ‘What is Local Optimality in Nonconvex-Nonconcave Minimax Optimization?’. arXiv
1902
Earlier work this paper cites.
1902
Earlier work this paper cites.
1903
Earlier work this paper cites.
1903
Earlier work this paper cites.
1903
Earlier work this paper cites.
1903
Earlier work this paper cites.
1905
Earlier work this paper cites.
1905
Earlier work this paper cites.
Kong, W. and R. D. Monteiro (2019), ‘An accelerated inexact proximal point method for solving nonconvex-concave min-max problems’. arXiv
1905
Earlier work this paper cites.
1905
Earlier work this paper cites.
1906
Earlier work this paper cites.
1906
Earlier work this paper cites.
1906
Earlier work this paper cites.
Jia, Z., L. F. Yang, and M. Wang (2019), ‘Feature-Based Q-Learning for Two-Player Stochastic Games’. arXiv
1906
Earlier work this paper cites.
1906
Earlier work this paper cites.
1906
Earlier work this paper cites.
1906
Earlier work this paper cites.
1906
Earlier work this paper cites.
1907
Earlier work this paper cites.
1907
Earlier work this paper cites.
1907
Earlier work this paper cites.
Weiss, P. (1907), ‘L’hypothèse du champ moléculaire et la propriété ferromagnétique’
1907
Earlier work this paper cites.
1908
Earlier work this paper cites.
1908
Earlier work this paper cites.
1909
Earlier work this paper cites.
1909
Earlier work this paper cites.
1909
Earlier work this paper cites.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
Bu, J., L. J. Ratliff, and M. Mesbahi (2019), ‘Global Convergence of Policy Gradient for Sequential Zero-Sum Linear Quadratic Dynamic Games’. arXiv
1911
Earlier work this paper cites.
1911
Earlier work this paper cites.
1911
Earlier work this paper cites.
1912
Earlier work this paper cites.
1912
Earlier work this paper cites.
1912
Earlier work this paper cites.
1912
Earlier work this paper cites.
1912
Earlier work this paper cites.
1912
Earlier work this paper cites.
1912
Earlier work this paper cites.
1912
Earlier work this paper cites.
Zermelo, E. and E. Borel (1913), ‘On an application of set theory to the theory of the game of chess’. In: Congress of Mathematicians
1913
Earlier work this paper cites.
Neumann, J. v. (1928), ‘Zur theorie der gesellschaftsspiele’. Mathematische annalen
1928
Earlier work this paper cites.
Von Neumann, J. and O. Morgenstern (1945), Theory of games and economic behavior
1945
Earlier work this paper cites.
Brown, G. W. (1951), ‘Iterative solution of games by fictitious play’. Activity analysis of production and allocation
1951
Earlier work this paper cites.
Dantzig, G. (1951), ‘A proof of the equivalence of the programming problem and the game problem, in “Activity analysis of production and allocation”(ed. TC Koopmans), Cowles Commission Monograph, No. 13’
1951
Earlier work this paper cites.
Nash, J. (1951), ‘Non-cooperative games’. Annals of mathematics
1951
Earlier work this paper cites.
Bellman, R. (1952), ‘On the theory of dynamic programming’. Proceedings of the National Academy of Sciences of the United States of America
1952
Earlier work this paper cites.
Shapley, L. S. (1953), ‘Stochastic games’. Proceedings of the national academy of sciences
1953
Earlier work this paper cites.
Minsky, M. L. (1954), Theory of neural-analog reinforcement systems and its application to the brain model problem
1954
Earlier work this paper cites.
Blackwell, D. et al. (1956), ‘An analog of the minimax theorem for vector payoffs.’. Pacific Journal of Mathematics
1956
Earlier work this paper cites.
Hannan, J. (1957), ‘Approximation to Bayes risk in repeated play’. Contributions to the Theory of Games
1957
Earlier work this paper cites.
Minsky, M. (1961), ‘Steps toward artificial intelligence’. Proceedings of the IRE
1961
Earlier work this paper cites.
Lemke, C. E. and J. T. Howson, Jr (1964), ‘Equilibrium points of bimatrix games’. Journal of the Society for Industrial and Applied Mathematics
1964
Earlier work this paper cites.
Selten, R. (1965), ‘Spieltheoretische behandlung eines oligopolmodells mit nachfrageträgheit: Teil i: Bestimmung des dynamischen preisgleichgewichts’. Zeitschrift für die gesamte Staatswissenschaft/Journal of Institutional and Theoretical Economics
1965
Earlier work this paper cites.
McKean, H. P. (1967), ‘Propagation of chaos for a class of non-linear parabolic equations’. Stochastic Differential Equations (Lecture Series in Differential Equations, Session 7, Catholic Univ., 1967)
1967
Earlier work this paper cites.
Martinet, B. (1970), ‘Regularisation, d’inéquations variationelles par approximations succesives’. Revue Francaise d’informatique et de Recherche operationelle
1970
Earlier work this paper cites.
Klopf, A. H. (1972), Brain function and adaptive systems: a heterostatic theory
1972
Earlier work this paper cites.
Maynard Smith, J. (1972), ‘On evolution’
1972
Earlier work this paper cites.
Simon, H. A. (1972), ‘Theories of bounded rationality’. Decision and organization
1972
Earlier work this paper cites.
Keynes, J. M. (1936), The General Theory of Employment, Interest and Money
1973
Earlier work this paper cites.
Smith, J. M. and G. R. Price (1973), ‘The logic of animal conflict’. Nature
1973
Earlier work this paper cites.
Shapley, L. S. (1974), ‘A note on the Lemke-Howson algorithm’. In: Pivoting and Extension
1974
Earlier work this paper cites.
Korpelevich, G. M. (1976), ‘The extragradient method for finding saddle points and other problems’. Matecon
1976
Earlier work this paper cites.
Rockafellar, R. T. (1976), ‘Augmented Lagrangians and applications of the proximal point algorithm in convex programming’. Mathematics of operations research
1976
Earlier work this paper cites.
Brown, N. and T. Sandholm (2015a), ‘Regret-Based Pruning in Extensive-Form Games.’. In: NIPS
1980
Earlier work this paper cites.
Kreps, D. M. and R. Wilson (1982), ‘Reputation and imperfect information’. Journal of economic theory
1982
Earlier work this paper cites.
Nemirovsky, A. S. and D. B. Yudin (1983), ‘Problem complexity and method efficiency in optimization.’
1983
Earlier work this paper cites.
Breton, M., J. A. Filar, A. Haurle, and T. A. Schultz (1986), ‘On the computation of equilibria in discounted stochastic dynamic games’. In: Dynamic games and applications in economics
1986
Earlier work this paper cites.
Papadimitriou, C. H. and J. N. Tsitsiklis (1987), ‘The complexity of Markov decision processes’. Mathematics of operations research
1987
Earlier work this paper cites.
Gärtner, J. (1988), ‘On the McKean-Vlasov limit for interacting diffusions’. Mathematische Nachrichten
1988
Earlier work this paper cites.
Jovanovic, B. and R. W. Rosenthal (1988), ‘Anonymous sequential games’. Journal of Mathematical Economics
1988
Earlier work this paper cites.
Sutton, R. S. (1988), ‘Learning to predict by the methods of temporal differences’. Machine learning
1988
Earlier work this paper cites.
Sznitman, A.-S. (1991), ‘Topics in propagation of chaos’. In: Ecole d’été de probabilités de Saint-Flour XIX—1989
1989
Earlier work this paper cites.
Koller, D. and N. Megiddo (1992), ‘The complexity of two-person zero-sum games in extensive form’. Games and economic behavior
1992
Earlier work this paper cites.
Lin, L.-J. (1992), ‘Self-improving reactive agents based on reinforcement learning, planning and teaching’. Machine learning
1992
Earlier work this paper cites.
Watkins, C. J. and P. Dayan (1992), ‘Q-learning’. Machine learning
1992
Earlier work this paper cites.
Williams, R. J. (1992), ‘Simple statistical gradient-following algorithms for connectionist reinforcement learning’. Machine learning
1992
Earlier work this paper cites.
Blume, L. E. (1993), ‘The statistical mechanics of strategic interaction’. Games and economic behavior
1993
Earlier work this paper cites.
Fudenberg, D. and D. M. Kreps (1993), ‘Learning mixed equilibria’. Games and economic behavior
1993
Earlier work this paper cites.
Tan, M. (1993), ‘Multi-agent reinforcement learning: Independent vs. cooperative agents’. In: Proceedings of the tenth international conference on machine learning
1993
Earlier work this paper cites.
Young, H. P. (1993), ‘The evolution of conventions’. Econometrica: Journal of the Econometric Society
1993
Earlier work this paper cites.
Littlestone, N. and M. K. Warmuth (1994), ‘The weighted majority algorithm’. Information and computation
1994
Earlier work this paper cites.
Littman, M. L. (1994), ‘Markov games as a framework for multi-agent reinforcement learning’. In: Machine learning proceedings 1994
1994
Earlier work this paper cites.
Osborne, M. J. and A. Rubinstein (1994), A course in game theory
1994
Earlier work this paper cites.
Auer, P., N. Cesa-Bianchi, Y. Freund, and R. E. Schapire (1995), ‘Gambling in a rigged casino: The adversarial multi-armed bandit problem’. In: Proceedings of IEEE 36th Annual Foundations of Computer Science
1995
Earlier work this paper cites.
Fudenberg, D. and D. Levine (1995), ‘Consistency and cautious fictitious play’. Journal of Economic Dynamics and Control
1995
Earlier work this paper cites.
Kaniovski, Y. M. and H. P. Young (1995), ‘Learning dynamics in games with stochastic perturbations’. Games and economic behavior
1995
Earlier work this paper cites.
Tesauro, G. (1995), ‘Temporal difference learning and TD-Gammon’. Communications of the ACM
1995
Earlier work this paper cites.
Bertsekas, D. P. and J. N. Tsitsiklis (1996), Neuro-dynamic programming
1996
Earlier work this paper cites.
Kaelbling, L. P., M. L. Littman, and A. W. Moore (1996), ‘Reinforcement learning: A survey’. Journal of artificial intelligence research
1996
Earlier work this paper cites.
Koller, D. and N. Megiddo (1996), ‘Finding mixed strategies with small supports in extensive form games’. International Journal of Game Theory
1996
Earlier work this paper cites.
Mahadevan, S. (1996), ‘Average reward reinforcement learning: Foundations, algorithms, and empirical results’. Machine learning
1996
Earlier work this paper cites.
Monderer, D. and L. S. Shapley (1996), ‘Potential games’. Games and economic behavior
1996
Earlier work this paper cites.
Sandholm, T. and R. Crites (1996), ‘Multiagent Reinforcement Learning in the Iterated Prisoner’s Dilemma’. Biosystems
1996
Earlier work this paper cites.
Schaeffer, J., R. Lake, P. Lu, and M. Bryant (1996), ‘Chinook the world man-machine checkers champion’. Ai Magazine
1996
Earlier work this paper cites.
Borkar, V. S. (1997), ‘Stochastic approximation with two time scales’. Systems & Control Letters
1997
Earlier work this paper cites.
Freund, Y. and R. E. Schapire (1997), ‘A decision-theoretic generalization of on-line learning and an application to boosting’. Journal of computer and system sciences
1997
Earlier work this paper cites.
von Stengel, B. and D. Koller (1997), ‘Team-maxmin equilibria’. Games and Economic Behavior
1997
Earlier work this paper cites.
Claus, C. and C. Boutilier (1998a), ‘The dynamics of reinforcement learning in cooperative multiagent systems’. AAAI/IAAI
1998
Earlier work this paper cites.
Claus, C. and C. Boutilier (1998b), ‘The dynamics of reinforcement learning in cooperative multiagent systems’. AAAI/IAAI
1998
Earlier work this paper cites.
Erev, I. and A. E. Roth (1998), ‘Predicting how people play games: Reinforcement learning in experimental games with unique, mixed strategy equilibria’. American economic review
1998
Earlier work this paper cites.
Fudenberg, D., F. Drew, D. K. Levine, and D. K. Levine (1998), The theory of learning in games
1998
Earlier work this paper cites.
Hu, J., M. P. Wellman, et al. (1998), ‘Multiagent reinforcement learning: theoretical framework and an algorithm.’. In: ICML
1998
Earlier work this paper cites.
Sutton, R. S. and A. G. Barto (1998), Reinforcement learning: An introduction
1998
Earlier work this paper cites.
Benaım, M. and M. W. Hirsch (1999), ‘Mixed equilibria and dynamical systems arising from fictitious play in perturbed games’. Games and Economic Behavior
1999
Earlier work this paper cites.
Boutilier, C., T. Dean, and S. Hanks (1999), ‘Decision-theoretic planning: Structural assumptions and computational leverage’. Journal of Artificial Intelligence Research
1999
Earlier work this paper cites.
Hu, J. (1999), Learning in dynamic noncooperative multiagent systems
1999
Earlier work this paper cites.
Jordan, M. I., Z. Ghahramani, T. S. Jaakkola, and L. K. Saul (1999), ‘An introduction to variational methods for graphical models’. Machine learning
1999
Earlier work this paper cites.
Weiss, G. (1999), Multiagent systems: a modern approach to distributed artificial intelligence
1999
Earlier work this paper cites.
Bowling, M. (2000), ‘Convergence problems of general-sum multiagent reinforcement learning’. In: ICML
2000
Earlier work this paper cites.
Bowling, M. and M. Veloso (2000), ‘An analysis of stochastic game theory for multiagent reinforcement learning’. Technical report, Carnegie-Mellon Univ Pittsburgh Pa School of Computer Science
2000
Earlier work this paper cites.
Konda, V. R. and J. N. Tsitsiklis (2000), ‘Actor-critic algorithms’. In: Advances in neural information processing systems
2000
Earlier work this paper cites.
Lauer, M. and M. Riedmiller (2000), ‘An algorithm for distributed reinforcement learning in cooperative multi-agent systems’. In: ICML
2000
Earlier work this paper cites.
Singh, S. P., M. J. Kearns, and Y. Mansour (2000), ‘Nash Convergence of Gradient Dynamics in General-Sum Games.’. In: UAI
2000
Earlier work this paper cites.
Stone, P. and M. Veloso (2000), ‘Multiagent systems: A survey from a machine learning perspective’. Autonomous Robots
2000
Earlier work this paper cites.
Sutton, R. S., D. A. McAllester, S. P. Singh, and Y. Mansour (2000), ‘Policy gradient methods for reinforcement learning with function approximation’. In: Advances in neural information processing systems
2000
Earlier work this paper cites.
Bowling, M. and M. Veloso (2001), ‘Rational and convergent learning in stochastic games’. In: International joint conference on artificial intelligence
2001
Earlier work this paper cites.
Hart, S. and A. Mas-Colell (2001), ‘A reinforcement procedure leading to correlated equilibrium’. In: Economics Essays
2001
Earlier work this paper cites.
Maskin, E. and J. Tirole (2001), ‘Markov perfect equilibrium: I. Observable actions’. Journal of Economic Theory
2001
Earlier work this paper cites.
Weaver, L. and N. Tao (2001), ‘The optimal reward baseline for gradient-based reinforcement learning’. In: Proceedings of the Seventeenth conference on Uncertainty in artificial intelligence
2001
Earlier work this paper cites.
Adler, J. L. and V. J. Blue (2002), ‘A cooperative multi-agent transportation management and route guidance system’. Transportation Research Part C: Emerging Technologies
2002
Earlier work this paper cites.
Auer, P., N. Cesa-Bianchi, Y. Freund, and R. E. Schapire (2002), ‘The nonstochastic multiarmed bandit problem’. SIAM journal on computing
2002
Earlier work this paper cites.
Ballard, D. and S. Zhu (2002), ‘Overcoming Non-Stationarity in Uncommunicative Learning’
2002
Earlier work this paper cites.
Bernstein, D. S., R. Givan, N. Immerman, and S. Zilberstein (2002), ‘The complexity of decentralized control of Markov decision processes’. Mathematics of operations research
2002
Earlier work this paper cites.
Borkar, V. S. (2002), ‘Reinforcement learning in Markovian evolutionary games’. Advances in Complex Systems
2002
Earlier work this paper cites.
Bowling, M. and M. Veloso (2002), ‘Multiagent learning using a variable learning rate’. Artificial Intelligence
2002
Earlier work this paper cites.
Brafman, R. I. and M. Tennenholtz (2002), ‘R-max-a general polynomial time algorithm for near-optimal reinforcement learning’. Journal of Machine Learning Research
2002
Earlier work this paper cites.
Camerer, C. F., T.-H. Ho, and J.-K. Chong (2002), ‘Sophisticated experience-weighted attraction learning and strategic teaching in repeated games’. Journal of Economic theory
2002
Earlier work this paper cites.
Campbell, M., A. J. Hoane Jr, and F.-h. Hsu (2002), ‘Deep blue’. Artificial intelligence
2002
Earlier work this paper cites.
Gigerenzer, G. and R. Selten (2002), Bounded rationality: The adaptive toolbox
2002
Earlier work this paper cites.
2002
Earlier work this paper cites.
Hofbauer, J. and W. H. Sandholm (2002), ‘On the global convergence of stochastic fictitious play’. Econometrica
2002
Earlier work this paper cites.
2002
Earlier work this paper cites.
2003
Earlier work this paper cites.
Beal, M. J. (2003), Variational algorithms for approximate Bayesian inference
2003
Earlier work this paper cites.
Billings, D., N. Burch, A. Davidson, R. Holte, J. Schaeffer, T. Schauenberg, and D. Szafron (2003), ‘Approximating game-theoretic optimal strategies for full-scale poker’. In: IJCAI
2003
Earlier work this paper cites.
Conitzer, V. and T. Sandholm (2003), ‘BL-WoLF: A framework for loss-bounded learnability in zero-sum games’. In: Proceedings of the 20th International Conference on Machine Learning (ICML-03)
2003
Earlier work this paper cites.
Even-Dar, E. and Y. Mansour (2003), ‘Learning rates for Q-learning’. Journal of machine learning Research
2003
Earlier work this paper cites.
Greenwald, A., K. Hall, and R. Serrano (2003), ‘Correlated Q-learning’. In: ICML
2003
Earlier work this paper cites.
2003
Earlier work this paper cites.
Hansen, N., S. D. Müller, and P. Koumoutsakos (2003), ‘Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (CMA-ES)’. Evolutionary computation
2003
Earlier work this paper cites.
Hu, J. and M. P. Wellman (2003), ‘Nash Q-learning for general-sum stochastic games’. Journal of machine learning research
2003
Earlier work this paper cites.
Lagoudakis, M. G. and R. Parr (2003), ‘Learning in zero-sum team markov games using factored value functions’. In: Advances in Neural Information Processing Systems
2003
Earlier work this paper cites.
Leslie, D. S., E. Collins, et al. (2003), ‘Convergent multiple-timescales reinforcement learning algorithms in normal form games’. The Annals of Applied Probability
2003
Earlier work this paper cites.
2003
Earlier work this paper cites.
Mannor, S. and N. Shimkin (2003), ‘The empirical Bayes envelope and regret minimization in competitive Markov decision processes’. Mathematics of Operations Research
2003
Earlier work this paper cites.
McMahan, H. B., G. J. Gordon, and A. Blum (2003), ‘Planning in the presence of cost functions controlled by an adversary’. In: Proceedings of the 20th International Conference on Machine Learning (ICML-03)
2003
Earlier work this paper cites.
Tuyls, K., K. Verbeeck, and T. Lenaerts (2003), ‘A selection-mutation model for q-learning in multi-agent systems’. In: Proceedings of the second international joint conference on Autonomous agents and multiagent systems
2003
Earlier work this paper cites.
Zinkevich, M. (2003), ‘Online convex programming and generalized infinitesimal gradient ascent’. In: Proceedings of the 20th international conference on machine learning (icml-03)
2003
Earlier work this paper cites.
Camerer, C. F., T.-H. Ho, and J.-K. Chong (2004), ‘A cognitive hierarchy model of games’. The Quarterly Journal of Economics
2004
Earlier work this paper cites.
Herings, P. J.-J., R. J. Peeters, et al. (2004), ‘Stationary equilibria in stochastic games: Structure, selection, and computation’. Journal of Economic Theory
2004
Earlier work this paper cites.
Kok, J. R. and N. Vlassis (2004), ‘Sparse cooperative Q-learning’. In: Proceedings of the twenty-first international conference on Machine learning
2004
Earlier work this paper cites.
Nemirovski, A. (2004), ‘Prox-method with rate of convergence O (1/t) for variational inequalities with Lipschitz continuous monotone operators and smooth convex-concave saddle point problems’. SIAM Journal on Optimization
2004
Earlier work this paper cites.
Tennenholtz, M. (2004), ‘Program equilibrium’. Games and Economic Behavior
2004
Earlier work this paper cites.
Bertsekas, D. P. (2005), ‘The dynamic programming algorithm’. Dynamic Programming and Optimal Control; Athena Scientific: Nashua, NH, USA
2005
Earlier work this paper cites.
Bowling, M. (2005), ‘Convergence and no-regret in multiagent learning’. In: Advances in neural information processing systems
2005
Earlier work this paper cites.
Daskalakis, C. and C. H. Papadimitriou (2005), ‘Three-player games are hard’. In: Electronic colloquium on computational complexity
2005
Earlier work this paper cites.
Even-Dar, E., S. M. Kakade, and Y. Mansour (2005), ‘Experts in a Markov decision process’. In: Advances in neural information processing systems
2005
Earlier work this paper cites.
Jan’t Hoen, P., K. Tuyls, L. Panait, S. Luke, and J. A. La Poutre (2005), ‘An overview of cooperative and competitive multiagent learning’. In: International Workshop on Learning and Adaption in Multi-Agent Systems
2005
Earlier work this paper cites.
Leslie, D. S. and E. J. Collins (2005), ‘Individual Q-learning in normal form games’. SIAM Journal on Control and Optimization
2005
Earlier work this paper cites.
Leyton-Brown, K. and M. Tennenholtz (2005), ‘Local-effect games’. In: Dagstuhl Seminar Proceedings
2005
Earlier work this paper cites.
Mguni, D. (2020), ‘Stochastic Potential Games’. arXiv preprint arXiv:2005.13527
2005
Earlier work this paper cites.
Morimoto, J. and K. Doya (2005), ‘Robust reinforcement learning’. Neural computation
2005
Earlier work this paper cites.
2005
Earlier work this paper cites.
Panait, L. and S. Luke (2005), ‘Cooperative multi-agent learning: The state of the art’. AAMAS
2005
Earlier work this paper cites.
Riedmiller, M. (2005), ‘Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method’. In: ECML
2005
Earlier work this paper cites.
Szer, D., F. Charpillet, and S. Zilberstein (2005), ‘MAA*: A Heuristic Search Algorithm for Solving Decentralized POMDPs’
2005
Earlier work this paper cites.
Tuyls, K. and A. Nowé (2005), ‘Evolutionary game theory and multi-agent reinforcement learning’
2005
Earlier work this paper cites.
Ye, Y. (2005), ‘A new complexity result on solving the Markov decision problem’. Mathematics of Operations Research
2005
Earlier work this paper cites.
Cesa-Bianchi, N. and G. Lugosi (2006), Prediction, learning, and games
2006
Earlier work this paper cites.
Chen, X. and X. Deng (2006), ‘Settling the complexity of two-player Nash equilibrium’. In: 2006 47th Annual IEEE Symposium on Foundations of Computer Science (FOCS’06)
2006
Cited alongside, same era.
2006
Cited alongside, same era.
Huang, M., R. P. Malhame, P. E. Caines, et al. (2006), ‘Large population stochastic dynamic games: closed-loop McKean-Vlasov systems and the Nash certainty equivalence principle’. Communications in Information & Systems
2006
Cited alongside, same era.
Kennedy, J. (2006), ‘Swarm intelligence’. In: Handbook of nature-inspired and innovative computing
2006
Cited alongside, same era.
Kocsis, L. and C. Szepesvári (2006), ‘Bandit based monte-carlo planning’. In: European conference on machine learning
Li, Y. (2017), ‘Deep reinforcement learning: An overview’. arXiv preprint arXiv:1701.07274
2017
Later among the works it cites.
Lowe, R., Y. Wu, A. Tamar, J. Harb, O. P. Abbeel, and I. Mordatch (2017), ‘Multi-agent actor-critic for mixed cooperative-competitive environments’. In: Advances in Neural Information Processing Systems
2017
Later among the works it cites.
Mescheder, L., S. Nowozin, and A. Geiger (2017), ‘The numerics of gans’. In: Advances in Neural Information Processing Systems
2017
Later among the works it cites.
Moravcik, M., M. Schmid, N. Burch, V. Lisy, D. Morrill, N. Bard, T. Davis, K. Waugh, M. Johanson, and M. Bowling (2017), ‘DeepStack: Expert-level artificial intelligence in heads-up no-limit poker’. Science
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2006
Cited alongside, same era.
Leslie, D. S. and E. J. Collins (2006), ‘Generalised weakened fictitious play’. Games and Economic Behavior
2006
Cited alongside, same era.
Li, L., T. J. Walsh, and M. L. Littman (2006), ‘Towards a unified theory of state abstraction for MDPs.’. AI&M
2006
Cited alongside, same era.
2006
Cited alongside, same era.
2006
Cited alongside, same era.
2006
Cited alongside, same era.
2006
Cited alongside, same era.
2006
Cited alongside, same era.
2017
Later among the works it cites.
Nedic, A., A. Olshevsky, and W. Shi (2017), ‘Achieving geometric convergence for distributed optimization over time-varying graphs’. SIAM Journal on Optimization
2017
Later among the works it cites.
Neu, G., A. Jonsson, and V. Gómez (2017), ‘A unified view of entropy-regularized markov decision processes’. NIPS
2017
Later among the works it cites.
Pham, H. and X. Wei (2017), ‘Dynamic programming for optimal control of stochastic McKean–Vlasov dynamics’. SIAM Journal on Control and Optimization
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Wei, C.-Y., Y.-T. Hong, and C.-J. Lu (2017), ‘Online reinforcement learning in stochastic games’. In: Advances in Neural Information Processing Systems
2017
Later among the works it cites.
Bailey, J. P. and G. Piliouras (2018), ‘Multiplicative weights update in zero-sum games’. In: Proceedings of the 2018 ACM Conference on Economics and Computation
2018
Later among the works it cites.
Brown, N. and T. Sandholm (2018), ‘Superhuman AI for heads-up no-limit poker: Libratus beats top professionals’. Science
2018
Later among the works it cites.
Brown, N., T. Sandholm, and B. Amos (2018), ‘Depth-limited solving for imperfect-information games’. Advances in neural information processing systems
2018
Later among the works it cites.
Cardaliaguet, P. and C.-A. Lehalle (2018), ‘Mean field game of controls and an application to trade crowding’. Mathematics and Financial Economics
2018
Later among the works it cites.
Carmona, R., F. Delarue, et al. (2018), Probabilistic Theory of Mean Field Games with Applications I-II
2018
Later among the works it cites.
Celli, A. and N. Gatti (2018), ‘Computational Results for Extensive-Form Adversarial Team Games’. In: AAAI Conference on Artificial Intelligence (AAAI)
2018
Later among the works it cites.
Dibangoye, J. and O. Buffet (2018), ‘Learning to Act in Decentralized Partially Observable MDPs’. In: International Conference on Machine Learning
2018
Later among the works it cites.
Farina, G., A. Celli, N. Gatti, and T. Sandholm (2018), ‘Ex ante coordination and collusion in zero-sum multi-player extensive-form games’. In: Advances in Neural Information Processing Systems
2018
Later among the works it cites.
Fazel, M., R. Ge, S. Kakade, and M. Mesbahi (2018), ‘Global Convergence of Policy Gradient Methods for the Linear Quadratic Regulator’. In: International Conference on Machine Learning
2018
Later among the works it cites.
Foerster, J. N., G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson (2018b), ‘Counterfactual Multi-Agent Policy Gradients’. In: S. A. McIlraith and K. Q. Weinberger (eds.): Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, Louisiana, USA, February 2-7, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
Haarnoja, T., A. Zhou, P. Abbeel, and S. Levine (2018), ‘Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor’. In: International Conference on Machine Learning
2018
Later among the works it cites.
Kroer, C. and T. Sandholm (2018), ‘A Unified Framework for Extensive-Form Game Abstraction with Bounds’. In: AI 3 Workshop at IJCAI
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
Li, Z. and A. Tewari (2018), ‘Sampled fictitious play is Hannan consistent’. Games and Economic Behavior
2018
Later among the works it cites.
2018
Later among the works it cites.
Macua, S. V., J. Zazo, and S. Zazo (2018), ‘Learning Parametric Closed-Loop Policies for Markov Potential Games’. In: International Conference on Learning Representations
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
Pachocki, J., G. Brockman, J. Raiman, S. Zhang, H. Pondé, J. Tang, F. Wolski, C. Dennison, R. Jozefowicz, P. Debiak, et al. (2018), ‘OpenAI Five, 2018’. URL https://blog. openai. com/openai-five
2018
Later among the works it cites.
Perolat, J., B. Piot, and O. Pietquin (2018), ‘Actor-critic fictitious play in simultaneous move multistage games’. In: International Conference on Artificial Intelligence and Statistics
2018
Later among the works it cites.
Pham, H. and X. Wei (2018), ‘Bellman equation and viscosity solutions for mean-field stochastic control problem’. ESAIM: Control, Optimisation and Calculus of Variations
2018
Later among the works it cites.
Rafique, H., M. Liu, Q. Lin, and T. Yang (2018), ‘Non-Convex Min-Max Optimization: Provable Algorithms and Applications in Machine Learning’. arXiv
2018
Later among the works it cites.
2018
Later among the works it cites.
Rothfuss, J., D. Lee, I. Clavera, T. Asfour, and P. Abbeel (2018), ‘ProMP: Proximal Meta-Policy Search’. In: International Conference on Learning Representations
2018
Later among the works it cites.
Saldi, N., T. Basar, and M. Raginsky (2018), ‘Markov–Nash Equilibria in Mean-Field Games with Discounted Cost’. SIAM Journal on Control and Optimization
2018
Later among the works it cites.
Sidford, A., M. Wang, X. Wu, L. Yang, and Y. Ye (2018), ‘Near-optimal time and sample complexities for solving Markov decision processes with a generative model’. In: Advances in Neural Information Processing Systems
2018
Later among the works it cites.
Silver, D., T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al. (2018), ‘A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play’. Science
2018
Later among the works it cites.
Song, M., A. Montanari, and P. Nguyen (2018), ‘A mean field view of the landscape of two-layers neural networks’. Proceedings of the National Academy of Sciences
2018
Later among the works it cites.
Srinivasan, S., M. Lanctot, V. Zambaldi, J. Pérolat, K. Tuyls, R. Munos, and M. Bowling (2018), ‘Actor-critic policy optimization in partially observable multiagent environments’. In: Advances in neural information processing systems
2018
Later among the works it cites.
Wei, H., G. Zheng, H. Yao, and Z. Li (2018), ‘Intellilight: A reinforcement learning approach for intelligent traffic light control’. In: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
2018
Later among the works it cites.
Wen, Y., Y. Yang, R. Luo, J. Wang, and W. Pan (2018), ‘Probabilistic Recursive Reasoning for Multi-Agent Reinforcement Learning’. In: International Conference on Learning Representations
2018
Later among the works it cites.
Zhang, K., Z. Yang, and T. Basar (2018a), ‘Networked multi-agent reinforcement learning in continuous spaces’. In: 2018 IEEE CDC
2018
Later among the works it cites.
Adolphs, L., H. Daneshmand, A. Lucchi, and T. Hofmann (2019), ‘Local saddle point optimization: A curvature exploitation approach’. In: The 22nd International Conference on Artificial Intelligence and Statistics
2019
Later among the works it cites.
Brown, N., A. Lerer, S. Gross, and T. Sandholm (2019), ‘Deep counterfactual regret minimization’. In: International Conference on Machine Learning
2019
Later among the works it cites.
Brown, N. and T. Sandholm (2019), ‘Superhuman AI for multiplayer poker’. Science
2019
Later among the works it cites.
Cecchin, A., P. D. Pra, M. Fischer, and G. Pelino (2019), ‘On the convergence problem in mean field games: a two state model without uniqueness’. SIAM Journal on Control and Optimization
2019
Later among the works it cites.
Cheung, Y. K. and G. Piliouras (2019), ‘Vortices instead of equilibria in minmax optimization: Chaos and butterfly effects of online learning in zero-sum games’. In: Conference on Learning Theory
2019
Later among the works it cites.
Da Silva, F. L. and A. H. R. Costa (2019), ‘A survey on transfer learning for multiagent reinforcement learning systems’. Journal of Artificial Intelligence Research
2019
Later among the works it cites.
Derakhshan, F. and S. Yousefi (2019), ‘A review on the applications of multiagent systems in wireless sensor networks’. International Journal of Distributed Sensor Networks
2019
Later among the works it cites.
Guo, X., A. Hu, R. Xu, and J. Zhang (2019), ‘Learning mean-field games’. In: Advances in Neural Information Processing Systems
2019
Later among the works it cites.
Hadikhanloo, S. and F. J. Silva (2019), ‘Finite mean field games: fictitious play and convergence to a first order continuous mean field game’. Journal de Mathématiques Pures et Appliquées
2019
Later among the works it cites.
Hernandez-Leal, P., B. Kartal, and M. E. Taylor (2019), ‘A survey and critique of multiagent deep reinforcement learning’. Autonomous Agents and Multi-Agent Systems
2019
Later among the works it cites.
Jaderberg, M., W. M. Czarnecki, I. Dunning, L. Marris, G. Lever, A. G. Castaneda, C. Beattie, N. C. Rabinowitz, A. S. Morcos, A. Ruderman, et al. (2019), ‘Human-level performance in 3D multiplayer games with population-based reinforcement learning’. Science
2019
Later among the works it cites.
Lacker, D. and T. Zariphopoulou (2019), ‘Mean field and n-agent games for optimal investment under relative performance criteria’. Mathematical Finance
2019
Later among the works it cites.
Liang, T. and J. Stokes (2019), ‘Interaction matters: A note on non-asymptotic local convergence of generative adversarial networks’. In: The 22nd International Conference on Artificial Intelligence and Statistics
2019
Later among the works it cites.
Mertikopoulos, P. and Z. Zhou (2019), ‘Learning in games with continuous action sets and unknown payoff functions’. Mathematical Programming
2019
Later among the works it cites.
Rosenberg, A. and Y. Mansour (2019), ‘Online Convex Optimization in Adversarial Markov Decision Processes’. In: International Conference on Machine Learning
2019
Later among the works it cites.
Saldi, N., T. Başar, and M. Raginsky (2019), ‘Approximate Nash equilibria in partially observed stochastic games with mean-field interactions’. Mathematics of Operations Research
2019
Later among the works it cites.
Schaefer, F. and A. Anandkumar (2019), ‘Competitive Gradient Descent’. In: H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (eds.): Advances in Neural Information Processing Systems
2019
Later among the works it cites.
Schmid, M., N. Burch, M. Lanctot, M. Moravcik, R. Kadlec, and M. Bowling (2019), ‘Variance reduction in monte carlo counterfactual regret minimization (VR-MCCFR) for extensive form games using baselines’. In: Proceedings of the AAAI Conference on Artificial Intelligence
2019
Later among the works it cites.
Shi, W., S. Song, and C. Wu (2019), ‘Soft policy gradient method for maximum entropy deep reinforcement learning’. In: Proceedings of the 28th International Joint Conference on Artificial Intelligence
2019
Later among the works it cites.
Son, K., D. Kim, W. J. Kang, D. E. Hostallero, and Y. Yi (2019), ‘Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning’. In: International Conference on Machine Learning
2019
Later among the works it cites.
Subramanian, J. and A. Mahajan (2019), ‘Reinforcement learning in stationary mean-field games’. In: Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems
2019
Later among the works it cites.
Swenson, B. and H. V. Poor (2019), ‘Smooth Fictitious Play in N × \times 2 Potential Games’. In: 2019 53rd Asilomar Conference on Signals, Systems, and Computers
2019
Later among the works it cites.
Tessler, C., Y. Efroni, and S. Mannor (2019), ‘Action robust reinforcement learning and applications in continuous control’. In: International Conference on Machine Learning
2019
Later among the works it cites.
Thekumparampil, K. K., P. Jain, P. Netrapalli, and S. Oh (2019), ‘Efficient algorithms for smooth minimax optimization’. In: Advances in Neural Information Processing Systems
2019
Later among the works it cites.
Vlatakis-Gkaragkounis, E.-V., L. Flokas, and G. Piliouras (2019), ‘Poincaré Recurrence, Cycles and Spurious Equilibria in Gradient-Descent-Ascent for Non-Convex Non-Concave Zero-Sum Games’. In: H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (eds.): Advances in Neural Information Processing Systems
2019
Later among the works it cites.
Wang, L., Q. Cai, Z. Yang, and Z. Wang (2019), ‘Neural Policy Gradient Methods: Global Optimality and Rates of Convergence’. In: International Conference on Learning Representations
2019
Later among the works it cites.
Wen, Y., Y. Yang, R. Luo, and J. Wang (2019), ‘Modelling Bounded Rationality in Multi-Agent Interactions by Generalized Recursive Reasoning’. IJCAI
2019
Later among the works it cites.
Zhang, Y. and M. M. Zavlanos (2019), ‘Distributed off-policy actor-critic reinforcement learning with policy consensus’. In: 2019 IEEE 58th Conference on Decision and Control (CDC)
2019
Later among the works it cites.
Zhou, M., Y. Chen, Y. Wen, Y. Yang, Y. Su, W. Zhang, D. Zhang, and J. Wang (2019), ‘Factorized Q-learning for large-scale multi-agent systems’. In: Proceedings of the First International Conference on Distributed Artificial Intelligence
2019
Later among the works it cites.
Bai, Y. and C. Jin (2020), ‘Provable self-play algorithms for competitive reinforcement learning’. In: International conference on machine learning
2020
Closest in time.
Bai, Y., C. Jin, and T. Yu (2020), ‘Near-optimal reinforcement learning with self-play’. Advances in neural information processing systems
2020
Closest in time.
Brown, N., A. Bakhtin, A. Lerer, and Q. Gong (2020), ‘Combining deep reinforcement learning and search for imperfect-information games’. Advances in Neural Information Processing Systems
2020
Closest in time.
Cheung, W. C., D. Simchi-Levi, and R. Zhu (2020), ‘Reinforcement learning for non-stationary markov decision processes: The blessing of (more) optimism’. ICML
2020
Closest in time.
Daskalakis, C., D. J. Foster, and N. Golowich (2020), ‘Independent policy gradient methods for competitive reinforcement learning’. Advances in neural information processing systems
2020
Closest in time.
Delarue, F. and R. F. Tchuendom (2020), ‘Selection of equilibria in a linear quadratic mean-field game’. Stochastic Processes and their Applications
2020
Closest in time.
Domingo-Enrich, C., S. Jelassi, A. Mensch, G. Rotskoff, and J. Bruna (2020), ‘A mean-field analysis of two-player zero-sum games’. Advances in neural information processing systems
2020
Closest in time.
Elie, R., J. Pérolat, M. Laurière, M. Geist, and O. Pietquin (2020), ‘On the Convergence of Model Free Learning in Mean Field Games.’. In: AAAI
2020
Closest in time.
Fiez, T., B. Chasnov, and L. Ratliff (2020), ‘Implicit learning dynamics in stackelberg games: Equilibria characterization, convergence analysis, and empirical study’. In: International Conference on Machine Learning
2020
Closest in time.
Golowich, N., S. Pattathil, C. Daskalakis, and A. Ozdaglar (2020), ‘Last iterate is slower than averaged iterate in smooth convex-concave saddle point problems’. In: Conference on Learning Theory
2020
Closest in time.
Haydari, A. and Y. Yılmaz (2020), ‘Deep reinforcement learning for intelligent transportation systems: A survey’. IEEE Transactions on Intelligent Transportation Systems
2020
Closest in time.
Ibrahim, A., W. Azizian, G. Gidel, and I. Mitliagkas (2020), ‘Linear lower bounds and conditioning of differentiable games’. In: International conference on machine learning
2020
Closest in time.
Letcher, A. (2018), ‘Stability and exploitation in differentiable games’. Ph.D. thesis, Master’s thesis, University of Oxford. Accessed: 2020-06-23
2020
Closest in time.
Lin, T., C. Jin, and M. Jordan (2020), ‘On gradient descent ascent for nonconvex-concave minimax problems’. In: International Conference on Machine Learning
2020
Closest in time.
Loizou, N., H. Berard, A. Jolicoeur-Martineau, P. Vincent, S. Lacoste-Julien, and I. Mitliagkas (2020), ‘Stochastic hamiltonian gradient methods for smooth games’. In: International Conference on Machine Learning
2020
Closest in time.
Michael, D. (2020), ‘Algorithmic Game Theory Lecture Notes’. http://www.cs.jhu.edu/ mdinitz/classes/AGT/Spring2020/Lectures/lecture6.pdf
2020
Closest in time.
Nguyen, T. T., N. D. Nguyen, and S. Nahavandi (2020), ‘Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications’. IEEE transactions on cybernetics
2020
Closest in time.
Nutz, M., J. San Martin, X. Tan, et al. (2020), ‘Convergence to the mean field game limit: a case study’. The Annals of Applied Probability
2020
Closest in time.
Ortner, R., P. Gajane, and P. Auer (2020), ‘Variational regret bounds for reinforcement learning’. In: Uncertainty in Artificial Intelligence
2020
Closest in time.
Schrittwieser, J., I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al. (2020), ‘Mastering atari, go, chess and shogi by planning with a learned model’. Nature
2020
Closest in time.
Sidford, A., M. Wang, L. Yang, and Y. Ye (2020), ‘Solving discounted stochastic two-player games with near-optimal time and sample complexity’. In: International Conference on Artificial Intelligence and Statistics
2020
Closest in time.
Sirignano, J. and K. Spiliopoulos (2020), ‘Mean field analysis of neural networks: A law of large numbers’. SIAM Journal on Applied Mathematics
2020
Closest in time.
uz Zaman, M. A., K. Zhang, E. Miehling, and T. Başar (2020), ‘Approximate equilibrium computation for discrete-time linear-quadratic mean-field games’. In: 2020 American Control Conference (ACC)
2020
Closest in time.
Xie, Q., Y. Chen, Z. Wang, and Z. Yang (2020), ‘Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium’. In: Conference on learning theory
2020
Closest in time.
Yang, Y., Y. Wen, L. Chen, J. Wang, K. Shao, D. Mguni, and W. Zhang (2020), ‘Multi-Agent Determinantal Q-Learning’
2020
Closest in time.
Zhang, B. and T. Sandholm (2020), ‘Small Nash equilibrium certificates in very large games’. Advances in Neural Information Processing Systems
2020
Closest in time.
Barazandeh, B., D. A. Tarzanagh, and G. Michailidis (2021), ‘Solving a class of non-convex min-max games using adaptive momentum methods’. In: ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
2021
Closest in time.
Cen, S., Y. Wei, and Y. Chi (2021), ‘Fast policy extragradient methods for competitive games with entropy regularization’. Advances in Neural Information Processing Systems
2021
Closest in time.
2021
Closest in time.
Daskalakis, C., M. Fishelson, and N. Golowich (2021), ‘Near-optimal no-regret learning in general games’. Advances in Neural Information Processing Systems
2021
Closest in time.
Deng, Y. and M. Mahdavi (2021), ‘Local stochastic gradient descent ascent: Convergence analysis and communication efficiency’. In: International Conference on Artificial Intelligence and Statistics
2021
Closest in time.
Domingues, O. D., P. Ménard, E. Kaufmann, and M. Valko (2021), ‘Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited’. In: Algorithmic Learning Theory
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
Franci, B. and S. Grammatico (2021), ‘Training generative adversarial networks via stochastic Nash games’. IEEE Transactions on Neural Networks and Learning Systems
2021
Closest in time.
Fu, H., W. Liu, S. Wu, Y. Wang, T. Yang, K. Li, J. Xing, B. Li, B. Ma, Q. Fu, et al. (2021), ‘Actor-Critic Policy Optimization in a Large-Scale Imperfect-Information Game’. In: International Conference on Learning Representations
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
He, H., S. Zhao, Y. Xi, and J. Ho (2021), ‘AGE: Enhancing the Convergence on GANs using Alternating extra-gradient with Gradient Extrapolation’. In: NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications
2021
Closest in time.
2021
Closest in time.
Kozuno, T., P. Ménard, R. Munos, and M. Valko (2021), ‘Learning in two-player zero-sum partially observable Markov games with perfect recall’. Advances in Neural Information Processing Systems
2021
Closest in time.
2021
Closest in time.
Li, H. (2021), ‘On the Complexity of Nonconvex-Strongly-Concave Smooth Minimax Optimization Using First-Order Methods’. Ph.D. thesis, Massachusetts Institute of Technology
2021
Closest in time.
2021
Closest in time.
Liu, Q., T. Yu, Y. Bai, and C. Jin (2021), ‘A sharp analysis of model-based reinforcement learning with self-play’. In: International Conference on Machine Learning
2021
Closest in time.
Loizou, N., H. Berard, G. Gidel, I. Mitliagkas, and S. Lacoste-Julien (2021), ‘Stochastic gradient descent-ascent and consensus optimization for smooth games: Convergence analysis under expected co-coercivity’. Advances in Neural Information Processing Systems
2021
Closest in time.
McAleer, S., J. Lanier, P. Baldi, and R. Fox (2021), ‘XDO: A double oracle algorithm for extensive-form games’. Advances in Neural Information Processing Systems (NeurIPS)
2021
Closest in time.
2021
Closest in time.
Mladenovic, A., I. Sakos, G. Gidel, and G. Piliouras (2021), ‘Generalized Natural Gradient Flows in Hidden Convex-Concave Games and GANs’. In: International Conference on Learning Representations
2021
Closest in time.
Perolat, J., R. Munos, J.-B. Lespiau, S. Omidshafiei, M. Rowland, P. Ortega, N. Burch, T. Anthony, D. Balduzzi, B. De Vylder, et al. (2021), ‘From Poincaré recurrence to convergence in imperfect information games: Finding equilibrium via regularization’. In: International Conference on Machine Learning
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
Anagnostides, I. and I. Panageas (2022), ‘Frequency-Domain Representation of First-Order Methods: A Simple and Robust Framework of Analysis’. In: Symposium on Simplicity in Algorithms (SOSA)
2022
Closest in time.
Bai, Y., C. Jin, S. Mei, and T. Yu (2022), ‘Near-optimal learning of extensive-form games with imperfect information’. In: International Conference on Machine Learning
2022
Closest in time.
2022
Closest in time.
BoHERE!HERE!hm, A., M. Sedlmayer, E. R. Csetnek, and R. I. Bot (2022), ‘Two Steps at a Time—Taking GAN Training in Stride with Tseng’s Method’. SIAM Journal on Mathematics of Data Science
2022
Closest in time.
Degrave, J., F. Felici, J. Buchli, M. Neunert, B. Tracey, F. Carpanese, T. Ewalds, R. Hafner, A. Abdolmaleki, D. de Las Casas, et al. (2022), ‘Magnetic control of tokamak plasmas through deep reinforcement learning’. Nature
2022
Closest in time.
2022
Closest in time.
Doan, T. (2022), ‘Convergence Rates of Two-Time-Scale Gradient Descent-Ascent Dynamics for Solving Nonconvex Min-Max Problems’. In: Learning for Dynamics and Control Conference
2022
Closest in time.
Du, Y., C. Ma, Y. Liu, R. Lin, H. Dong, J. Wang, and Y. Yang (2022b), ‘Scalable model-based policy optimization for decentralized networked systems’. In: 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
2022
Closest in time.
(FAIR)†, M. F. A. R. D. T., A. Bakhtin, N. Brown, E. Dinan, G. Farina, C. Flaherty, D. Fried, A. Goff, J. Gray, H. Hu, et al. (2022), ‘Human-level play in the game of Diplomacy by combining language models with strategic reasoning’. Science
2022
Closest in time.
Farina, G. and T. Sandholm (2022), ‘Fast payoff matrix sparsification techniques for structured extensive-form games’. In: Proceedings of the AAAI Conference on Artificial Intelligence
2022
Closest in time.
Gorbunov, E., N. Loizou, and G. Gidel (2022), ‘Extragradient method: O (1/K) last-iterate convergence for monotone variational inequalities and connections with cocoercivity’. In: International Conference on Artificial Intelligence and Statistics
2022
Closest in time.
Ha, J. and G. Kim (2022), ‘On Convergence of Lookahead in Smooth Games’. In: International Conference on Artificial Intelligence and Statistics
2022
Closest in time.
2022
Closest in time.
He, H., S. Zhao, Y. Xi, J. Ho, and Y. Saad (2022), ‘GDA-AM: ON THE EFFECTIVENESS OF SOLVING MIN-IMAX OPTIMIZATION VIA ANDERSON MIXING’. In: International Conference on Learning Representations
2022
Closest in time.
Mao, W. and T. Başar (2022), ‘Provably efficient reinforcement learning in decentralized general-sum markov games’. Dynamic Games and Applications
2022
Closest in time.
Perolat, J., B. De Vylder, D. Hennes, E. Tarassov, F. Strub, V. de Boer, P. Muller, J. T. Connor, N. Burch, T. Anthony, et al. (2022), ‘Mastering the game of Stratego with model-free multiagent reinforcement learning’. Science
2022
Closest in time.
Qu, G., A. Wierman, and N. Li (2022), ‘Scalable reinforcement learning for multiagent networked systems’. Operations Research
2022
Closest in time.
Sharma, P., R. Panda, G. Joshi, and P. Varshney (2022), ‘Federated minimax optimization: Improved convergence analyses and algorithms’. In: International Conference on Machine Learning
2022
Closest in time.
2022
Closest in time.
Song, Z., S. Mei, and Y. Bai (2022), ‘Sample-efficient learning of correlated equilibria in extensive-form games’. Advances in Neural Information Processing Systems
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
Yang, J., A. Orvieto, A. Lucchi, and N. He (2022), ‘Faster single-loop algorithms for minimax optimization without strong concavity’. In: International Conference on Artificial Intelligence and Statistics
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
Zhang, B. and T. Sandholm (2022), ‘Polynomial-time optimal equilibria with a mediator in extensive-form games’. Advances in Neural Information Processing Systems
2022
Closest in time.
Fiegel, C., P. Ménard, T. Kozuno, R. Munos, V. Perchet, and M. Valko (2023), ‘Adapting to game trees in zero-sum imperfect information games’. In: International Conference on Machine Learning
2023
Closest in time.
Liu, W., H. Fu, Q. Fu, and Y. Wei (2023), ‘Opponent-limited online search for imperfect information games’. In: International Conference on Machine Learning
2023
Closest in time.
2023
Closest in time.
McAleer, S., G. Farina, G. Zhou, M. Wang, Y. Yang, and T. Sandholm (2023), ‘Team-PSRO for learning approximate TMECor in large team games via cooperative reinforcement learning’. Advances in Neural Information Processing Systems
2023
Closest in time.
2023
Closest in time.
Tang, X., S. M. McAleer, Y. Yang, et al. (2023), ‘Regret-minimizing double oracle for extensive-form games’. In: International Conference on Machine Learning
2023
Closest in time.
2023
Closest in time.
Anagnostides, I., I. Panageas, G. Farina, and T. Sandholm (2024), ‘On the convergence of no-regret learning dynamics in time-varying games’. Advances in Neural Information Processing Systems
2024
Closest in time.
2024
Closest in time.
Ge, Z., Z. Xu, T. Ding, W. Li, and Y. Gao (2024), ‘Efficient subgame refinement for extensive-form games’. Advances in Neural Information Processing Systems
2024
Closest in time.
Ma, C., A. Li, Y. Du, H. Dong, and Y. Yang (2024), ‘Efficient and scalable reinforcement learning for large-scale network control’. Nature Machine Intelligence
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Xu, L., D. Perez-Liebana, and A. Dockhorn (2024b), ‘Strategy Game-Playing with Size-Constrained State Abstraction’. In: 2024 IEEE Conference on Games (CoG)
2024
Closest in time.
2024
Closest in time.
Szepesvári, C. and M. L. Littman (1999), ‘A unified analysis of value-function-based reinforcement-learning algorithms’. Neural computation
2060
Closest in time.
Qu, G., Y. Lin, A. Wierman, and N. Li (2020), ‘Scalable multi-agent reinforcement learning for networked systems with average reward’. Advances in Neural Information Processing Systems
2086
Closest in time.
Sunehag, P., G. Lever, A. Gruslys, W. M. Czarnecki, V. Zambaldi, M. Jaderberg, M. Lanctot, N. Sonnerat, J. Z. Leibo, K. Tuyls, et al. (2018), ‘Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward’. In: Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems
2087
Closest in time.