Fetching the paper…
Reading the bibliography…
In this article, we review recent Deep Learning advances in the context of how they have been applied to play different types of video games such as first-person shooters, arcade games, and real-time strategy games.
A general framework for parallel distributed processing
D. E. Rumelhart, G. E. Hinton, J. L. McClelland, et al · 1986
Earlier work this paper cites.
Modèles connexionnistes de l’apprentissage
Y. Le Cun · 1987
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
R. J. Williams and D. Zipser · 1989
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
B. Hassibi, D. G. Stork, et al · 1993
Earlier work this paper cites.
Reinforcement learning for robots using neural networks
L.-J. Lin · 1993
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
M. Tan · 1993
Earlier work this paper cites.
On-line Q-learning using connectionist systems
G. A. Rummery and M. Niranjan · 1994
Earlier work this paper cites.
Hq-learning
M. Wiering and J. Schmidhuber · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, Y. Mansour, et al · 1999
Earlier work this paper cites.
Overview of robocup-98
M. Asada, M. M. Veloso, M. Tambe, I. Noda, H. Kitano, and G. K. Kraetzschmar · 2000
Earlier work this paper cites.
TORCS, the open racing car simulator
B. Wymann, E. Espié, C. Guionneau, C. Dimitrakakis, R. Coulom, and A. Sumner · 2000
Earlier work this paper cites.
Keepaway soccer: A machine learning test bed
P. Stone and R. S. Sutton · 2001
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
A. G. Barto and S. Mahadevan · 2003
Earlier work this paper cites.
All learning is local: Multi-agent learning in global reward games
Y.-H. Chang, T. Ho, and L. P. Kaelbling · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
N. Chentanez, A. G. Barto, and S. P. Singh · 2005
Earlier work this paper cites.
Real-time neuroevolution in the NERO video game
K. O. Stanley, B. D. Bryant, and R. Miikkulainen · 2005
Earlier work this paper cites.
Evolutionary computation and games
S. M. Lucas and G. Kendall · 2006
Earlier work this paper cites.
Computational intelligence in games
R. Miikkulainen, B. D. Bryant, R. Cornelius, I. V. Karpov, K. O. Stanley, and C. H. Yong · 2006
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
S. Legg and M. Hutter · 2007
Earlier work this paper cites.
Machine learning in digital games: a survey
L. Galway, D. Charles, and M. Black · 2008
Earlier work this paper cites.
Exploiting open-endedness to solve problems through the search for novelty
J. Lehman and K. O. Stanley · 2008
Earlier work this paper cites.
Emergence in games
P. Sweetser · 2008
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
Automatic content generation in the galactic arms race video game
E. J. Hastings, R. K. Guha, and K. O. Stanley · 2009
Earlier work this paper cites.
Ontogenetic and phylogenetic reinforcement learning
J. Togelius, T. Schaul, D. Wierstra, C. Igel, F. Gomez, and J. Schmidhuber · 2009
Earlier work this paper cites.
Double q-learning
H. V. Hasselt · 2010
Earlier work this paper cites.
A new design for a turing test for bots
P. Hingston · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
J. Schmidhuber · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
J. S. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl · 2011
Earlier work this paper cites.
Measuring intelligence through games
T. Schaul, J. Togelius, and J. Schmidhuber · 2011
Earlier work this paper cites.
UT⁁ 2: Human-like behavior via neuroevolution of combat behavior and replay of human traces
J. Schrum, I. V. Karpov, and R. Miikkulainen · 2011
Earlier work this paper cites.
A survey of monte carlo tree search methods
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton · 2012
Earlier work this paper cites.
Model-free reinforcement learning with continuous action in practice
T. Degris, P. M. Pilarski, and R. S. Sutton · 2012
Earlier work this paper cites.
Believable Bots: Can Computers Play Like People?
P. Hingston · 2012
Earlier work this paper cites.
Neurovisual control in the quake ii environment
M. Parker and B. D. Bryant · 2012
Earlier work this paper cites.
Combining search-based procedural content generation and social gaming in the petalz video game
S. Risi, J. Lehman, D. B. D’Ambrosio, R. Hall, and K. O. Stanley · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures
J. Bergstra, D. Yamins, and D. D. Cox · 2013
Earlier work this paper cites.
Evolving large-scale neural networks for vision-based reinforcement learning
J. Koutník, G. Cuccu, J. Schmidhuber, and F. Gomez · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Learning and game AI
H. Muñoz-Avila, C. Bauckhage, M. Bida, C. B. Congdon, and G. Kendall · 2013
Earlier work this paper cites.
The combinatorial multi-armed bandit problem and its application to real-time strategy games
S. Ontanón · 2013
Earlier work this paper cites.
Imitating human playing styles in super mario bros
J. Ortega, N. Shaker, J. Togelius, and G. N. Yannakakis · 2013
Earlier work this paper cites.
A video game description language for model-based or interactive learning
T. Schaul · 2013
Earlier work this paper cites.
The turing test track of the 2012 mario AI championship: entries and evaluation
N. Shaker, J. Togelius, G. N. Yannakakis, L. Poovanna, V. S. Ethiraj, S. J. Johansson, R. G. Reynolds, L. K. Heether, T. Schumann, and M. Gallagher · 2013
Earlier work this paper cites.
Player modeling
G. N. Yannakakis, P. Spronck, D. Loiacono, and E. André · 2013
Earlier work this paper cites.
A. Graves, G. Wayne, and I. Danihelka · 2014
Earlier work this paper cites.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
X. Guo, S. Singh, H. Lee, R. L. Lewis, and X. Wang · 2014
Earlier work this paper cites.
A neuroevolution approach to general atari game playing
M. Hausknecht, J. Lehman, R. Miikkulainen, and P. Stone · 2014
Earlier work this paper cites.
Generative agents for player decision modeling in games
C. Holmgård, A. Liapis, J. Togelius, and G. N. Yannakakis · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2015
Earlier work this paper cites.
Deep apprenticeship learning for playing video games
M. Bogdanovic, D. Markovikj, M. Denil, and N. De Freitas · 2015
Cited alongside, same era.
Deepdriving: Learning affordance for direct perception in autonomous driving
C. Chen, A. Seff, A. Kornhauser, and J. Xiao · 2015
Cited alongside, same era.
Deep recurrent q-learning for partially observable mdps
M. Hausknecht and P. Stone · 2015
Cited alongside, same era.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Massively parallel methods for deep reinforcement learning
A. Nair, P. Srinivasan, S. Blackwell, C. Alcicek, R. Fearon, A. De Maria, V. Panneershelvam, M. Suleyman, C. Beattie, S. Petersen, et al · 2015
Pathnet: Evolution channels gradient descent in super neural networks
C. Fernando, D. Banarse, C. Blundell, Y. Zwols, D. Ha, A. A. Rusu, A. Pritzel, and D. Wierstra · 2017
Closest in time.
Stabilising experience replay for deep multi-agent reinforcement learning
J. Foerster, N. Nardelli, G. Farquhar, P. Torr, P. Kohli, S. Whiteson, et al · 2017
Closest in time.
What can you do with a rock? affordance extraction viaword embeddings
N. Fulda, D. Ricks, B. Murdoch, and D. Wingate · 2017
Closest in time.
Game engine learning from video
M. Guzdial, B. Li, and M. O. Riedl · 2017
Closest in time.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2017
Closest in time.
Evocommander: A novel game based on evolving and switching between artificial brains
D. Jallov, S. Risi, and J. Togelius · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language understanding for textbased games using deep reinforcement learning
K. Narasimhan, T. D. Kulkarni, and R. Barzilay · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh · 2015
Cited alongside, same era.
Neuroevolution in games: State of the art and open challenges
S. Risi and J. Togelius · 2015
Cited alongside, same era.
Deep learning in neural networks: An overview
J. Schmidhuber · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
A panorama of artificial and computational intelligence in games
G. N. Yannakakis and J. Togelius · 2015
Cited alongside, same era.
Closest in time.
Continual online evolution for in-game build order adaptation in starcraft
N. Justesen and S. Risi · 2017
Closest in time.
Learning macromanagement in StarCraft from replays using deep learning
N. Justesen and S. Risi · 2017
Closest in time.
Schema networks: Zero-shot transfer with a generative causal model of intuitive physics
K. Kansky, T. Silver, D. A. Mély, M. Eldawy, M. Lázaro-Gredilla, X. Lou, N. Dorfman, S. Sidor, S. Phoenix, and D. George · 2017
Closest in time.
Beating atari with natural language guided reinforcement learning
R. Kaplan, C. Sauer, and A. Sosa · 2017
Closest in time.
Multi-task learning in atari video games with emergent tangled program graphs
S. Kelly and M. I. Heywood · 2017
Closest in time.
Overcoming catastrophic forgetting in neural networks
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, et al · 2017
Closest in time.
Text-based adventures of the golovin ai agent
B. Kostka, J. Kwiecieli, J. Kowalski, and P. Rychlikowski · 2017
Closest in time.
General video game ai: Learning from screen capture
K. Kunanusont, S. M. Lucas, and D. Perez-Liebana · 2017
Closest in time.
Playing FPS games with deep reinforcement learning
G. Lample and D. S. Chaplot · 2017
Closest in time.
Multi-agent reinforcement learning in sequential social dilemmas
J. Z. Leibo, V. Zambaldi, M. Lanctot, J. Marecki, and T. Graepel · 2017
Closest in time.
Deep reinforcement learning: An overview
Y. Li · 2017
Closest in time.
Teacher-student curriculum learning
T. Matiisen, A. Oliver, T. Cohen, and J. Schulman · 2017
Closest in time.
Count-based exploration with neural density models
G. Ostrovski, M. G. Bellemare, A. van den Oord, and R. Munos · 2017
Closest in time.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Closest in time.
Multiagent bidirectionally-coordinated nets for learning to play starcraft combat games
P. Peng, Q. Yuan, Y. Wen, Y. Yang, Z. Tang, H. Long, and J. Wang · 2017
Closest in time.
DLNE: a hybridization of deep learning and neuroevolution for visual control
A. Precht, M. Thorhauge, M. Hvilshøj, and S. Risi · 2017
Closest in time.
Evolution strategies as a scalable alternative to reinforcement learning
T. Salimans, J. Ho, X. Chen, and I. Sutskever · 2017
Closest in time.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Closest in time.
F. P. Such, V. Madhavan, E. Conti, J. Lehman, K. O. Stanley, and J. Clune · 2017
Closest in time.
Multiagent cooperation and competition with deep reinforcement learning
A. Tampuu, T. Matiisen, D. Kodelja, I. Kuzovkin, K. Korjus, J. Aru, J. Aru, and R. Vicente · 2017
Closest in time.
Distral: Robust multitask reinforcement learning
Y. Teh, V. Bapst, R. Pascanu, N. Heess, J. Quan, J. Kirkpatrick, W. M. Czarnecki, and R. Hadsell · 2017
Closest in time.
A deep hierarchical approach to lifelong learning in minecraft
C. Tessler, S. Givony, T. Zahavy, D. J. Mankowitz, and S. Mannor · 2017
Closest in time.
Elf: An extensive, lightweight and flexible research platform for real-time strategy games
Y. Tian, Q. Gong, W. Shang, Y. Wu, and C. L. Zitnick · 2017
Closest in time.
Episodic exploration for deep deterministic policies: An application to starcraft micromanagement tasks
N. Usunier, G. Synnaeve, Z. Lin, and S. Chintala · 2017
Closest in time.
Hybrid reward architecture for reinforcement learning
H. Van Seijen, R. Laroche, M. Fatemi, and J. Romoff · 2017
Closest in time.
Starcraft II: A new challenge for reinforcement learning
O. Vinyals, T. Ewalds, S. Bartunov, A. S. Georgiev, Petko Vezhnevets, M. Yeo, A. Makhzani, H. Kuttler, J. Agapiou, J. Schrittwieser, S. Gaffney, S. Petersen, K. Simonyan, T. Schaul, H. v. Hasselt, D. Silver, T. Lillicrap, K. Calderone, P. Keet, A. Brunasso, D. Lawrence, A. Ekermo, J. Repp, and R. Tsing · 2017
Closest in time.
Sample efficient actor-critic with experience replay
Z. Wang, V. Bapst, N. Heess, V. Mnih, R. Munos, K. Kavukcuoglu, and N. de Freitas · 2017
Closest in time.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Y. Wu, E. Mansimov, R. B. Grosse, S. Liao, and J. Ba · 2017
Closest in time.
Training agent for first-person shooter game with actor-critic curriculum learning
Y. Wu and Y. Tian · 2017
Closest in time.
A deep compositional framework for human-like language acquisition in virtual environment
H. Yu, H. Zhang, and W. Xu · 2017
Closest in time.
Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents
E. Conti, V. Madhavan, F. P. Such, J. Lehman, K. Stanley, and J. Clune · 2018
Closest in time.
Textworld: A learning environment for text-based games
M.-A. Côté, A. Kádár, X. Yuan, B. Kybartas, T. Barnes, E. Fine, J. Moore, M. Hausknecht, L. E. Asri, M. Adada, W. Tay, and A. Trischler · 2018
Closest in time.
Playing atari with six neurons
G. Cuccu, J. Togelius, and P. Cudre-Mauroux · 2018
Closest in time.
IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Closest in time.
Counterfactual multi-agent policy gradients
J. Foerster, G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson · 2018
Closest in time.
Noisy networks for exploration
M. Fortunato, M. G. Azar, B. Piot, J. Menick, M. Hessel, I. Osband, A. Graves, V. Mnih, R. Munos, D. Hassabis, O. Pietquin, C. Blundell, and S. Legg · 2018
Closest in time.
Human-like playtesting with deep learning
S. Gudmundsson, P. Eisen, E. Poromaa, A. Nodet, S. Purmonen, R. Meurling, B. Kozakowski, and L. Cao · 2018
Closest in time.
D. Ha and J. Schmidhuber · 2018
Closest in time.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. G. Azar, and D. Silver · 2018
Closest in time.
Deep q-learning from demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, G. Dulac-Arnold, et al · 2018
Closest in time.
Distributed prioritized experience replay
D. Horgan, J. Quan, D. Budden, G. Barth-Maron, M. Hessel, H. Van Hasselt, and D. Silver · 2018
Closest in time.
Illuminating generalization in deep reinforcement learning through procedural level generation
N. Justesen, R. R. Torrado, P. Bontrager, A. Khalifa, J. Togelius, and S. Risi · 2018
Closest in time.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
M. C. Machado, M. G. Bellemare, E. Talvitie, J. Veness, M. Hausknecht, and M. Bowling · 2018
Closest in time.
Observe and look further: Achieving consistent performance on atari
T. Pohlen, B. Piot, T. Hester, M. G. Azar, D. Horgan, D. Budden, G. Barth-Maron, H. van Hasselt, J. Quan, M. Večerík, et al · 2018
Closest in time.
Deep reinforcement learning for general video game AI
R. Rodriguez Torrado, P. Bontrager, J. Togelius, J. Liu, and D. Perez-Liebana · 2018
Closest in time.
Procedural content generation via machine learning (pcgml)
A. Summerville, S. Snodgrass, M. J. Guzdial, C. Holmgx00E5rd, A. K. Hoover, A. Isaksen, A. Nealen, and J. Togelius · 2018
Closest in time.
Tstarbots: Defeating the cheating level builtin ai in starcraft ii in the full game
P. Sun, X. Sun, L. Han, J. Xiong, Q. Wang, B. Li, Y. Zheng, J. Liu, Y. Liu, H. Liu, et al · 2018
Closest in time.
Reinforcement learning for build-order production in starcraft ii
Z. Tang, D. Zhao, Y. Zhu, and P. Guo · 2018
Closest in time.
Evolving mario levels in the latent space of a deep convolutional generative adversarial network
V. Volz, J. Schrum, J. Liu, S. M. Lucas, A. Smith, and S. Risi · 2018
Closest in time.
Artificial Intelligence and Games
G. N. Yannakakis and J. Togelius · 2018
Closest in time.
Learn what not to learn: Action elimination with deep reinforcement learning
T. Zahavy, M. Haroush, N. Merlis, D. J. Mankowitz, and S. Mannor · 2018
Closest in time.