Fetching the paper…
Reading the bibliography…
Although recent work in AI has made great progress in solving large, zero-sum, extensive-form games, the underlying assumption in most past work is that the parameters of the game itself are known to the agents.
Quantal response equilibria for normal form games
Richard D McKelvey and Thomas R Palfrey · 1995
Earlier work this paper cites.
Efficient computation of behavior strategies
Bernhard Von Stengel · 1996
Earlier work this paper cites.
Quantal response equilibria for extensive form games
Richard D McKelvey and Thomas R Palfrey · 1998
Earlier work this paper cites.
An analysis of stochastic game theory for multiagent reinforcement learning
Michael Bowling and Manuela Veloso · 2000
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Learning payoff functions in infinite games
Yevgeniy Vorobeychik, Michael P Wellman, and Satinder Singh · 2007
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter · 2008
Earlier work this paper cites.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione · 2008
Earlier work this paper cites.
Learning and approximating the optimal strategy to commit to
Joshua Letchford, Vincent Conitzer, and Kamesh Munagala · 2009
Earlier work this paper cites.
Using game theory for los angeles airport security
James Pita, Manish Jain, Fernando Ordónez, Christopher Portway, Milind Tambe, Craig Western, Praveen Paruchuri, and Sarit Kraus · 2009
Cited alongside, same era.
Smoothing techniques for computing nash equilibria of sequential games
Samid Hoda, Andrew Gilpin, Javier Pena, and Tuomas Sandholm · 2010
Cited alongside, same era.
Computational rationalization: the inverse equilibrium problem
Kevin Waugh, Brian D Ziebart, and J Andrew Bagnell · 2011
Cited alongside, same era.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Learning optimal commitment to overcome insecurity
Avrim Blum, Nika Haghtalab, and Ariel D Procaccia · 2014
Cited alongside, same era.
Deploying paws: Field optimization of the protection assistant for wildlife security
Fei Fang, Thanh Hong Nguyen, Rob Pickles, Wai Y Lam, Gopalasamy R Clements, Bo An, Amandeep Singh, Milind Tambe, and Andrew Lemieux · 2016
Later among the works it cites.
Stephen Gould, Basura Fernando, Anoop Cherian, Peter Anderson, Rodrigo Santa Cruz, and Edison Guo · 2016
Later among the works it cites.
Composing graphical models with neural networks for structured representations and fast inference
Matthew Johnson, David K Duvenaud, Alex Wiltschko, Ryan P Adams, and Sandeep R Datta · 2016
Later among the works it cites.
Learning in games via reinforcement and regularization
Panayotis Mertikopoulos and William H Sandholm · 2016
Later among the works it cites.
Optnet: Differentiable optimization as a layer in neural networks
Brandon Amos and J Zico Kolter · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Heads-up limit hold’em poker is solved
Michael Bowling, Neil Burch, Michael Johanson, and Oskari Tammelin · 2015
Cited alongside, same era.
Learning equilibria of games via payoff queries
John Fearnley, Martin Gairing, Paul W Goldberg, and Rahul Savani · 2015
Cited alongside, same era.
Gradient methods for stackelberg security games
Kareem Amin, Satinder Singh, and Michael P Wellman · 2016
Cited alongside, same era.
Later among the works it cites.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2017
Later among the works it cites.
Theoretical and practical advances on smoothing for extensive-form games
Christian Kroer, Kevin Waugh, Fatma Kilinc-Karzan, and Tuomas Sandholm · 2017
Later among the works it cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisỳ, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.