Fetching the paper…
Reading the bibliography…
This paper introduces MazeBase: an environment for simple 2D games, designed as a sandbox for machine learning approaches to reasoning and planning.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 1992
Earlier work this paper cites.
Computer Go: an AI oriented survey
Bouzy, Bruno and Cazenave, Tristan · 2001
Earlier work this paper cites.
Brood War API, 2008-
BWAPI · 2008
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
The 2010 mario AI championship: Level generation track
Shaker, N., Togelius, J., Yannakakis, G., Weber, B., Shimizu, T., Hashiyama, N., Soreson, P., Pasquier, P, Mawhorter, G., Takahashi, G., Smith, R., and Baumgarten, R · 2010
Earlier work this paper cites.
Monte-carlo tree search in ms. pac-man
N., Ikehata and Ito, T · 2011
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Graves, Alex, Antonoglou, Ioannis, Wierstra, Daan, and Riedmiller, Martin · 2013
Cited alongside, same era.
Neural turing machines
Graves, Alex, Wayne, Greg, and Danihelka, Ivo · 2014
Cited alongside, same era.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Guo, Xiaoxiao, Singh, Satinder, Lee, Honglak, Lewis, Richard L, and Wang, Xiaoshi · 2014
Cited alongside, same era.
The GVG-AI competition
Perez, D., Samothrakis, S., Togelius, J., Schaul, T., and Lucas, S · 2014
Cited alongside, same era.
Inferring algorithmic patterns with stack-augmented recurrent nets
Joulin, Armand and Mikolov, Tomas · 2015
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A., Veness, Joel, Bellemare, Marc G., Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K., Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Closest in time.
Adaapt: A deep architecture for adaptive policy transfer from multiple sources
Rajendran, J., Prasanna, P., Ravindran, B., and Khapra, M · 2015
Closest in time.
End-to-end memory networks
Sukhbaatar, Sainbayar, Szlam, Arthur, Weston, Jason, and Fergus, Rob · 2015
Closest in time.
Pointer networks
Vinyals, Oriol, Fortunato, Meire, and Jaitly, Navdeep · 2015
Closest in time.
Reinforcement learning neural turing machines
Zaremba, Wojciech and Sutskever, Ilya · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A roadmap towards machine intelligence
Mikolov, Tomas, Joulin, Armand, and Baroni, Marco · 2015
Cited alongside, same era.
Towards ai-complete question answering: A set of prerequisite toy tasks
Weston, J., Bordes, A., Chopra, S., and Mikolov, T
Cited in the paper.
Memory networks
Weston, J., Chopra, S., and Bordes, A
Cited in the paper.
Zaremba, Wojciech, Mikolov, Tomas, Joulin, Armand, and Fergus, Rob · 2015
Closest in time.