Fetching the paper…
Reading the bibliography…
Recent progress in artificial intelligence through reinforcement learning (RL) has shown great success on increasingly complex single-agent environments and two-player turn-based games.
Neue begründung der theorie quadratischer formen von unendlichvielen veränderlichen
Ernst Hellinger · 1909
Earlier work this paper cites.
Iterative solutions of games by fictitious play
G. W. Brown · 1951
Earlier work this paper cites.
The Rating of Chessplayers, Past and Present
Arpad E Elo · 1978
Earlier work this paper cites.
Interactions between learning and evolution
David Ackley and Michael Littman · 1991
Earlier work this paper cites.
Learning complex, extended sequences using the principle of history compression
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
G. Tesauro · 1995
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
Salah El Hihi and Yoshua Bengio · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Robocup: A challenge problem for ai and robotics
Hiroaki Kitano, Minoru Asada, Yasuo Kuniyoshi, Itsuki Noda, Eiichi Osawai, and Hitoshi Matsubara · 1997
Earlier work this paper cites.
New methods for competitive coevolution
Christopher D Rosin and Richard K Belew · 1997
Earlier work this paper cites.
Quake III Arena, 1999
id Software · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Between MDPs and Semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder P. Singh · 1999
Earlier work this paper cites.
An introduction to collective intelligence
David H Wolpert and Kagan Tumer · 1999
Earlier work this paper cites.
The complexity of decentralized control of Markov Decision Processes
Daniel S. Bernstein, Shlomo Zilberstein, and Neil Immerman · 2000
Earlier work this paper cites.
Layered learning
Peter Stone and Manuela Veloso · 2000
Earlier work this paper cites.
Human-level ai’s killer application: Interactive computer games
John Laird and Michael VanLent · 2001
Earlier work this paper cites.
The Quake III Arena Bot (Master’s Thesis), 2001
J. M. P. Van Waveren · 2001
Earlier work this paper cites.
The effects of action video game experience on the time course of inhibition of return and the efficiency of visual search
Alan D Castel, Jay Pratt, and Emily Drummond · 2005
Earlier work this paper cites.
Fundamental components of the gameplay experience: Analysing immersion
Laura Ermi and Frans Mäyrä · 2005
Earlier work this paper cites.
Three states and a plan: the A.I. of F.E.A.R
Jeff Orkin · 2006
Earlier work this paper cites.
On experiences in a complex and competitive gaming domain: Reinforcement learning meets robocup
Martin Riedmiller and Thomas Gabel · 2007
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens J P Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Where do rewards come from?
Satinder Singh, Richard L Lewis, and Andrew G Barto · 2009
Earlier work this paper cites.
Hierarchical controller learning in a first-person shooter
Niels Van Hoorn, Julian Togelius, and Jurgen Schmidhuber · 2009
Earlier work this paper cites.
Learning model-free robot control by a Monte Carlo EM algorithm
Nikos Vlassis, Marc Toussaint, Georgios Kontes, and Savas Piperidis · 2009
Cited alongside, same era.
Intrinsically motivated reinforcement learning: An evolutionary perspective
Satinder Singh, Richard L Lewis, Andrew G Barto, and Jonathan Sorg · 2010
Cited alongside, same era.
Ad hoc autonomous agent teams: Collaboration without pre-coordination
Peter Stone, Gal A Kaminka, Sarit Kraus, Jeffrey S Rosenschein, et al · 2010
Cited alongside, same era.
On the limits of the human motor control precision: the search for a device’s human resolution
François Bérard, Guangyu Wang, and Jeremy R Cooperstock · 2011
Cited alongside, same era.
Independent reinforcement learners in cooperative Markov games: a survey regarding coordination problems
Laëtitia Matignon, Guillaume J. Laurent, and Nadine Le Fort-Piat · 2012
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Later among the works it cites.
Learning multiagent communication with backpropagation
Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus · 2016
Later among the works it cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Later among the works it cites.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio · 2017
Later among the works it cites.
Learning with opponent-learning awareness
Jakob N Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rodrigo Quian Quiroga · 2012
Cited alongside, same era.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Variational policy search via trajectory optimization
Sergey Levine and Vladlen Koltun · 2013
Cited alongside, same era.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Cited alongside, same era.
Jan Koutník, Klaus Greff, Faustino Gomez, and Jürgen Schmidhuber · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Cited alongside, same era.
Generating sentences from a continuous space
Samuel R Bowman, Luke Vilnis, Oriol Vinyals, Andrew M Dai, Rafal Jozefowicz, and Samy Bengio · 2015
Cited alongside, same era.
Neuroscience-inspired artificial intelligence
Demis Hassabis, Dharshan Kumaran, Christopher Summerfield, and Matthew Botvinick · 2017
Later among the works it cites.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et al · 2017
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z. Leibo, David Silver, and Koray Kavukcuoglu · 2017
Later among the works it cites.
Neuroscience needs behavior: correcting a reductionist bias
John W Krakauer, Asif A Ghazanfar, Alex Gomez-Marin, Malcolm A MacIver, and David Poeppel · 2017
Later among the works it cites.
Playing FPS games with deep reinforcement learning
Guillaume Lample and Devendra Singh Chaplot · 2017
Later among the works it cites.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinícius Flores Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolat, David Silver, and Thore Graepel · 2017
Later among the works it cites.
Multi-agent reinforcement learning in sequential social dilemmas
Joel Z Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Later among the works it cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravcik, Martin Schmid, Neil Burch, Viliam Lisy, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu · 2017
Later among the works it cites.
Reinforced variational inference
Theophane Weber and Nicolas Heess · 2017
Later among the works it cites.
Training agent for first-person shooter game with actor-critic curriculum learning
Yuxin Wu and Yuandong Tian · 2017
Later among the works it cites.
QuakeCon, 2018
2018
Closest in time.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2018
Closest in time.
Impala: Scalable distributed Deep-RL with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Closest in time.
UT Austin Villa: RoboCup 2017 3D simulation league competition and technical challenges champions
Patrick MacAlpine and Peter Stone · 2018
Closest in time.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel · 2018
Closest in time.