Fetching the paper…
Reading the bibliography…
Mastering a video game requires skill, tactics and strategy.
A world championship caliber checkers program
J. Schaeffer, J. Culberson, N. Treloar, B. Knight, P. Lu, and D. Szafron · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
G. Tesauro · 1995
Earlier work this paper cites.
Improving elevator performance using reinforcement learning
R. Crites and A. Barto · 1996
Earlier work this paper cites.
Reinforcement learning in the multi-robot domain
M. J. Matarić · 1997
Earlier work this paper cites.
Value-function reinforcement learning in markov games
M. L. Littman · 2001
Earlier work this paper cites.
Deep blue
M. Campbell, A. J. Hoane, and F.-h. Hsu · 2002
Earlier work this paper cites.
Super mario evolution
J. Togelius, S. Karakovskiy, J. Koutník, and J. Schmidhuber · 2009
Earlier work this paper cites.
Multi-agent reinforcement learning: An overview
L. Buşoniu, R. Babuška, and B. De Schutter · 2010
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Hierarchical reinforcement learning for robot navigation
B. Bischoff, D. Nguyen-Tuong, I.-H. Lee, F. Streichert, and A. Knoll · 2013
Cited alongside, same era.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
I. J. Goodfellow, M. Mirza, D. Xiao, A. Courville, and Y. Bengio · 2013
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Closest in time.
The malmo platform for artificial intelligence experimentation
M. Johnson, K. Hofmann, T. Hutton, and D. Bignell · 2016
Closest in time.
Libretro
libRetro site · 2016
Closest in time.
Long-term planning by short-term prediction
S. Shalev-Shwartz, N. Ben-Zrihem, A. Cohen, and A. Shashua · 2016
Closest in time.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Closest in time.
Universe
Universe · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Van Hasselt, A. Guez, and D. Silver · 2015
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Z. Wang, N. de Freitas, and M. Lanctot · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
M. G. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
Closest in time.
Target-driven visual navigation in indoor scenes using deep reinforcement learning
Y. Zhu, R. Mottaghi, E. Kolve, J. J. Lim, A. Gupta, L. Fei-Fei, and A. Farhadi · 2016
Closest in time.