Fetching the paper…
Reading the bibliography…
OpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games.
Open-ended learning in symmetric zero-sum games
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech M. Czarnecki, Julien Pérolat, Max Jaderberg, and Thore Graepel · 1901
Earlier work this paper cites.
The hanabi challenge: A new frontier for AI research
Nolan Bard, Jakob N. Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H. Francis Song, Emilio Parisotto, Vincent Dumoulin, Subhodeep Moitra, Edward Hughes, Iain Dunning, Shibl Mourad, Hugo Larochelle, Marc G. Bellemare, and Michael Bowling · 1902
Earlier work this paper cites.
Computing approximate equilibria in sequential adversarial games by exploitability descent
Edward Lockhart, Marc Lanctot, Julien Pérolat, Jean-Baptiste Lespiau, Dustin Morrill, Finbarr Timbers, and Karl Tuyls · 1903
Earlier work this paper cites.
Rethinking formal models of partially observable multiagent decision making
Vojtech Kovarík, Martin Schmid, Neil Burch, Michael Bowling, and Viliam Lisý · 1906
Earlier work this paper cites.
Shayegan Omidshafiei, Daniel Hennes, Dustin Morrill, Rémi Munos, Julien Pérolat, Marc Lanctot, Audrunas Gruslys, Jean-Baptiste Lespiau, and Karl Tuyls · 1906
Earlier work this paper cites.
Behaviour suite for reinforcement learning
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepezvari, Satinder Singh, Benjamin Van Roy, Richard Sutton, David Silver, and Hado Van Hasselt · 1908
Earlier work this paper cites.
Simplified two-person Poker
H. W. Kuhn · 1950
Earlier work this paper cites.
Iterative solutions of games by fictitious play
G. W. Brown · 1951
Earlier work this paper cites.
An iterative method of solving a game
J Robinson · 1951
Earlier work this paper cites.
Game-playing and game-learning automata
D. Michie · 1966
Earlier work this paper cites.
An analysis of alpha-beta pruning
Donald E. Knuth and Ronald W Moore · 1975
Earlier work this paper cites.
The *-minimax search procedure for trees containing chance nodes
B. W. Ballard · 1983
Earlier work this paper cites.
Three problems in learning mixed-strategy Nash equilibria
J. S. Jordan · 1993
Earlier work this paper cites.
Fast algorithms for finding randomized strategies in game trees
D. Koller, N. Megiddo, and B. von Stengel · 1994
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
A Course in Game Theory
M.J. Osborne and A. Rubinstein · 1994
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Y. Freund and R. E. Shapire · 1995
Earlier work this paper cites.
Evolutionary Games and Population Dynamics
Josef Hofbauer and Karl Sigmund · 1998
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
S. Hart and A. Mas-Colell · 2000
Earlier work this paper cites.
Nash convergence of gradient dynamics in general-sum games
Satinder P. Singh, Michael J. Kearns, and Yishay Mansour · 2000
Earlier work this paper cites.
Comparing policy-gradient algorithms, 2001
Richard S. Sutton, Satinder Singh, and David McAllester · 2001
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Michael Bowling and Manuela Veloso · 2002
Earlier work this paper cites.
Policy gradient methods for control applications
Jan Peters · 2002
Earlier work this paper cites.
Analyzing Complex Strategic Interactions in Multi-Agent Systems
William E Walsh, Rajarshi Das, Gerald Tesauro, and Jeffrey O Kephart · 2002
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
H. McMahan, G. Gordon, and A. Blum · 2003
Earlier work this paper cites.
Choosing samples to compute heuristic-strategy Nash equilibrium
W. E. Walsh, D. C. Parkes, and R. Das · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
Solving the oshi-zumo game
M. Buro · 2004
Earlier work this paper cites.
Optimal play of the dice game pig
Todd W. Neller and Clifton G.M. Presser · 2004
Earlier work this paper cites.
Convergence and no-regret in multiagent learning
Michael Bowling · 2005
Cited alongside, same era.
General game-playing: Overview of the AAAI competition
M. Genesereth, N. Love, and B. Pell · 2005
Cited alongside, same era.
Bayes’ bluff: Opponent modelling in poker
Finnegan Southey, Michael Bowling, Bryce Larson, Carmelo Piccione, Neil Burch, Darse Billings, and Chris Rayner · 2005
Cited alongside, same era.
Bandit-based Monte Carlo planning
L. Kocsis and C. Szepesvári · 2006
Cited alongside, same era.
Methods for empirical game-theoretic analysis
Michael P. Wellman · 2006
Cited alongside, same era.
Efficient selectivity and backup operators in Monte-Carlo tree search
R. Coulom · 2007
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
Solving heads-up limit Texas Hold’em
Oskari Tammelin, Neil Burch, Michael Johanson, and Michael Bowling · 2015
Later among the works it cites.
Solving games with functional regret estimation
Kevin Waugh, Dustin Morrill, J. Andrew Bagnell, and Michael Bowling · 2015
Later among the works it cites.
Tensorflow: A system for large-scale machine learning
Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Later among the works it cites.
Algorithms for computing strategies in two-player simultaneous move games
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A multiagent reinforcement learning algorithm with non-linear dynamics
Sherief Abdallah and Victor Lesser · 2008
Cited alongside, same era.
Regret minimization in games with incomplete information
M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione · 2008
Cited alongside, same era.
The complexity of computing a Nash equilibrium
Constantinos Daskalakis, Paul W Goldberg, and Christos H Papadimitriou · 2009
Cited alongside, same era.
Sampling for regret minimization in extensive games
M. Lanctot, K. Waugh, M. Bowling, and M. Zinkevich · 2009
Cited alongside, same era.
Artificial Intelligence: A Modern Approach
S. Russell and P. Norvig · 2009
Cited alongside, same era.
Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations
Y. Shoham and K. Leyton-Brown · 2009
Cited alongside, same era.
Branislav Bošanský, Viliam Lisý, Marc Lanctot, Jiří Čermák, and Mark H.M. Winands · 2016
Later among the works it cites.
Opponent modeling in deep reinforcement learning
He He, Jordan L. Boyd-Graber, Kevin Kwok, and Hal Daumé III · 2016
Later among the works it cites.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Using regret estimation to solve games compactly
Dustin Morrill · 2016
Later among the works it cites.
Softened approximate policy iteration for markov games
Julien Pérolat, Bilal Piot, Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2016
Later among the works it cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2017
Later among the works it cites.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Perolat, David Silver, and Thore Graepel · 2017
Later among the works it cites.
Multi-agent reinforcement learning in sequential social dilemmas
Joel Z. Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel · 2017
Later among the works it cites.
Deal or no deal? End-to-end learning of negotiation dialogues
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra · 2017
Later among the works it cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisý, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Later among the works it cites.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
David Balduzzi, Karl Tuyls, Julien Perolat, and Thore Graepel · 2018
Later among the works it cites.
Emergent communication through negotiation
Kris Cao, Angeliki Lazaridou, Marc Lanctot, Joel Z. Leibo, Karl Tuyls, and Stephen Clark · 2018
Later among the works it cites.
Bayesian action decoder for deep multi-agent reinforcement learning
Jakob N. Foerster, H. Francis Song, Edward Hughes, Neil Burch, Iain Dunning, Shimon Whiteson, Matthew Botvinick, and Michael Bowling · 2018
Later among the works it cites.
Fast deep reinforcement learning using online adjustments from the past
Steven Hansen, Pablo Sprechmann, Alexander Pritzel, André Barreto, and Charles Blundell · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew J. Hausknecht, and Michael Bowling · 2018
Later among the works it cites.
Modeling others using oneself in multi-agent reinforcement learning
Roberta Raileanu, Emily Denton, Arthur Szlam, and Rob Fergus · 2018
Later among the works it cites.
Actor-critic policy optimization in partially observable multiagent environments
Sriram Srinivasan, Marc Lanctot, Vinicius Zambaldi, Julien Perolat, Karl Tuyls, Remi Munos, and Michael Bowling · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
R. Sutton and A. Barto · 2018
Later among the works it cites.
A Generalised Method for Empirical Game Theoretic Analysis
Karl Tuyls, Julien Perolat, Marc Lanctot, Joel Z Leibo, and Thore Graepel · 2018
Later among the works it cites.
Deep counterfactual regret minimization
Noam Brown, Adam Lerer, Sam Gross, and Tuomas Sandholm · 2019
Closest in time.
Superhuman AI for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Closest in time.
Hex, the full story
Ryan B. Hayward and Bjarne Toft · 2019
Closest in time.
α \alpha -rank: Multi-agent evaluation by evolution
Shayegan Omidshafiei, Christos Papadimitriou, Georgios Piliouras, Karl Tuyls, Mark Rowland, Jean-Baptiste Lespiau, Wojciech M. Czarnecki, Marc Lanctot, Julien Perolat, and Remi Munos · 2019
Closest in time.