Fetching the paper…
Reading the bibliography…
Game-based decision-making involves reasoning over both world dynamics and strategic interactions among the agents.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
The linear complementarity problem
B. Curtis Eaves · 1971
Earlier work this paper cites.
Cognitron: A self-organizing multilayered neural network
Kunihiko Fukushima · 1975
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Ronald J. Williams and David Zipser · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S. Sutton · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S. Sutton · 1991
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman · 1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Analyzing complex strategic interactions in multi-agent systems
William Walsh, Rajarshi Das, Gerald Tesauro, and Jeffrey Kephart · 2002
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
H. Brendan McMahan, Geoffrey J Gordon, and Avrim Blum · 2003
Earlier work this paper cites.
A novel method for automatic strategy acquisition in N N -player non-zero-sum games
S. Phelps, M. Marcinkiewicz, and S. Parsons · 2006
Earlier work this paper cites.
Methods for empirical game-theoretic analysis
Michael P. Wellman · 2006
Earlier work this paper cites.
Combining opponent modeling and model-based reinforcement learning in a two-player competitive game
Brian Collins · 2007
Earlier work this paper cites.
Exploring large strategy spaces in empirical game modeling
L. Julian Schvartzman and Michael P. Wellman · 2009
Earlier work this paper cites.
Stronger CDA strategies through empirical game-theoretic analysis and reinforcement learning
L. Julian Schvartzman and Michael P. Wellman · 2009
Earlier work this paper cites.
Lab experiments for the study of social-ecological systems
Marco A. Janssen, Robert Holahan, Allen Lee, and Elinor Ostrom · 2010
Earlier work this paper cites.
Strategy exploration in empirical games
Patrick R. Jordan, L. Julian Schvartzman, and Michael P. Wellman · 2010
Earlier work this paper cites.
Probabilistic analysis of simulation-based games
Yevgeniy Vorobeychik · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Goeffrey J. Gordon, and J. Andrew Bagnell · 2011
Earlier work this paper cites.
Texplore: Real-time sample-efficient reinforcement learning for robots
Todd Hester and Peter Stone · 2012
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
Richard S Sutton, Csaba Szepesvári, Alborz Geramifard, and Michael P. Bowling · 2012
Earlier work this paper cites.
Online implicit agent modelling
Nolan Bard, Michael Johanson, Neil Burch, and Michael Bowling · 2013
Earlier work this paper cites.
Modeling deep temporal dependencies with recurrent grammar cells
Vicent Michalski, Roland Memisevic, and Kishore Konda · 2014
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
Stéphane Ross and J. Andrew Bagnell · 2014
Earlier work this paper cites.
Model regularization for stable sample rollouts
Erin Talvitie · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Opponent modeling by expectation–maximization and sequence prediction in simplified poker
Richard Mealing and Jonathan L Shapiro · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard Lewis, and Satinder Singh · 2015
Cited alongside, same era.
From pixels to torques: Policy learning with deep dynamical models
Niklas Wahlström, Thomas B. Schön, and Marc Peter Deisenroth · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Cited alongside, same era.
Opponent modeling in deep reinforcement learning
He He, Jordan Boyd-Graber, Kevin Kwok, and Hal Daumé III · 2016
Cited alongside, same era.
Gambit: Software tools for game theory
Richard D. McKelvey, Andrew M. McLennan, and Theodore L. Turocy · 2016
Cited alongside, same era.
DeepMDP: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G. Bellemare · 2019
Later among the works it cites.
https://github.com/HumanCompatibleAI/multi-agent , 2019
HumanCompatibleAI · 2019
Later among the works it cites.
Deep state-space models in multi-agent systems
Pararawendy Indarjo · 2019
Later among the works it cites.
OpenSpiel: A framework for reinforcement learning in games
Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, Satyaki Upadhyay, Julien Pérolat, Sriram Srinivasan, Finbarr Timbers, Karl Tuyls, Shayegan Omidshafiei, Daniel Hennes, Dustin Morrill, Paul Muller, Timo Ewalds, Ryan Faulkner, János Kramár, Bart De Vylder, Brennan Saeta, James Bradbury, David Ding, Sebastian Borgeaud, Matthew Lai, Julian Schrittwieser, Thomas Anthony, Edward Hughes, Ivo Danihelka, and Jonah Ryan-Davis · 2019
Later among the works it cites.
Stochastic prediction of multi-agent interactions from partial observations
Chen Sun, Per Karlsson, Jiajun Wu, Joshua B Tenenbaum, and Kevin Murphy · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Recurrent environment simulators
Silvia Chiappa, Sébastien Racaniere, Daan Wierstra, and Shakir Mohamed · 2017
Cited alongside, same era.
A survey of learning in multiagent environments: Dealing with non-stationarity
Pablo Hernandez-Leal, Michael Kaisers, Tim Baarslag, and Enrique Munoz de Cote · 2017
Cited alongside, same era.
Uncertainty-driven imagination for continuous deep reinforcement learning
Gabriel Kalweit and Joschka Boedecker · 2017
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrūnas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolat, David Silver, and Thore Graepel · 2017
Cited alongside, same era.
Multi-agent reinforcement learning in sequential social dilemmas
Joel Z. Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel · 2017
Cited alongside, same era.
Later among the works it cites.
A regularized opponent model with maximum entropy objective
Zheng Tian, Ying Wen, Zhichen Gong, Faiz Punakkath, Shihao Zou, and Jun Wang · 2019
Later among the works it cites.
Ready policy one: World building through active learning
Philip Ball, Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski, and Stephen Roberts · 2020
Later among the works it cites.
Haiku: Sonnet for JAX, 2020
Tom Hennigan, Trevor Cai, Tamara Norman, and Igor Babuschkin · 2020
Later among the works it cites.
Acme: A research framework for distributed reinforcement learning
Matthew W. Hoffman, Bobak Shahriari, John Aslanides, Gabriel Barth-Maron, Nikola Momchev, Danila Sinopalnikov, Piotr Stańczyk, Sabela Ramos, Anton Raichuk, Damien Vincent, Léonard Hussenot, Robert Dadashi, Gabriel Dulac-Arnold, Manu Orsini, Alexis Jacq, Johan Ferret, Nino Vieillard, Seyed Kamyar Seyed Ghasemipour, Sertan Girgin, Olivier Pietquin, Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, Sarah Henderson, Abe Friesen, Ruba Haroun, Alex Novikov, Sergio Gómez Colmenarejo, Serkan Cabi, Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Andrew Cowie, Ziyu Wang, Bilal Piot, and Nando de Freitas · 2020
Later among the works it cites.
Multi-agent reinforcement learning with multi-step generative models
Orr Krupnik, Igor Mordatch, and Aviv Tamar · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, and David Silver · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Later among the works it cites.
Learning to play against any mixture of opponents
Max Olan Smith, Thomas Anthony, Yongzhao Wang, and Michael P. Wellman · 2020
Later among the works it cites.
Bounds and dynamics for empirical game theoretic analysis
Karl Tuyls, Julien Pérolat, Marc Lanctot, Edward Hughes, Richard Everett, Joel Z. Leibo, Csaba Szepesvári, and Thore Graepel · 2020
Later among the works it cites.
Model-based reinforcement learning for decentralized multiagent rendezvous
Rose E Wang, Chase Kew, Dennis Lee, Edward Lee, Brian Andrew Ichter, Tingnan Zhang, Jie Tan, and Aleksandra Faust · 2020
Later among the works it cites.
MOPO: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Model-based multi-agent rl in zero-sum markov games with near-optimal sample complexity
Kaiqing Zhang, Sham Kakade, Tamer Basar, and Lin Yang · 2020
Later among the works it cites.
Reverb: A framework for experience replay, 2021
Albin Cassirer, Gabriel Barth-Maron, Eugene Brevdo, Sabela Ramos, Toby Boyd, Thibault Sottiaux, and Manuel Kroiss · 2021
Later among the works it cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2021
Later among the works it cites.
Scalable evaluation of multi-agent reinforcement learning with melting pot
Joel Z. Leibo, Edgar Duéñez-Guzmán, Alexander Sasha Vezhnevets, John P. Agapiou, Peter Sunehag, Raphael Koster, Jayd Matyas, Charles Beattie, Igor Mordatch, and Thore Graepel · 2021
Later among the works it cites.
XDO: A double oracle algorithm for extensive-form games
Stephen McAleer, John Lanier, Kevin Wang, Pierre Baldi, and Roy Fox · 2021
Later among the works it cites.
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Johan S. Obando-Ceron and Pablo Samuel Castro · 2021
Later among the works it cites.
Iterative empirical game solving via single policy best response
Max Olan Smith, Thomas Anthony, and Michael P. Wellman · 2021
Later among the works it cites.
MAMBPO: Sample-efficient multi-robot reinforcement learning using learned world models
Daniël Willemsen, Mario Coppola, and Guido CHE de Croon · 2021
Later among the works it cites.
Launchpad: A programming model for distributed machine learning research
Fan Yang, Gabriel Barth-Maron, Piotr Stańczyk, Matthew Hoffman, Siqi Liu, Manuel Kroiss, Aedan Pope, and Alban Rrustemi · 2021
Later among the works it cites.
Exploiting extensive-form structure in empirical game-theoretic analysis
Christine Konicki, Mithun Chakraborty, and Michael P. Wellman · 2022
Later among the works it cites.
Centralized model and exploration policy for multi-agent RL
Qizhen Zhang, Chris Lu, Animesh Garg, and Jakob Foerster · 2022
Later among the works it cites.
Coursera neural networks for machine learning lecture 6
Geoffrey Hinton · 2023
Closest in time.
Search-improved game-theoretic multiagent reinforcement learning in general and negotiation games (extended abstract)
Zun Li, Marc Lanctot, Kevin McKee, Luke Marris, Ian Gemp, Daniel Hennes, Paul Muller, Kate Larson, Yoram Bachrach, and Michael P. Wellman · 2023
Closest in time.