Fetching the paper…
Reading the bibliography…
In this paper, we introduce a new set of reinforcement learning (RL) tasks in Minecraft (a flexible 3D world).
Mazes, maps, and memory
Olton, David S · 1979
Earlier work this paper cites.
Q-learning
Watkins, Christopher JCH and Dayan, Peter · 1992
Earlier work this paper cites.
Temporal difference learning and td-gammon
Tesauro, Gerald · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
A robot that reinforcement-learns to identify and memorize important previous observations
Bakker, Bram, Zhumatiy, Viktor, Gruener, Gabriel, and Schmidhuber, Jürgen · 2003
Earlier work this paper cites.
Bandit based monte-carlo planning
Kocsis, Levente and Szepesvári, Csaba · 2006
Earlier work this paper cites.
Deep auto-encoder neural networks in reinforcement learning
Lange, Sascha and Riedmiller, Martin · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, Vinod and Hinton, Geoffrey E · 2010
Earlier work this paper cites.
Recurrent policy gradients
Wierstra, Daan, Förster, Alexander, Peters, Jan, and Schmidhuber, Jürgen · 2010
Earlier work this paper cites.
Torch7: A matlab-like environment for machine learning
Collobert, Ronan, Kavukcuoglu, Koray, and Farabet, Clément · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Graves, Alex · 2013
Cited alongside, same era.
Guided policy search
Levine, Sergey and Koltun, Vladlen · 2013
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, Ross, Donahue, Jeff, Darrell, Trevor, and Malik, Jitendra · 2014
Cited alongside, same era.
Graves, Alex, Wayne, Greg, and Danihelka, Ivo · 2014
Cited alongside, same era.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Guo, Xiaoxiao, Singh, Satinder, Lee, Honglak, Lewis, Richard L, and Wang, Xiaoshi · 2014
Cited alongside, same era.
Universal value function approximators
Schaul, Tom, Horgan, Daniel, Gregor, Karol, and Silver, David · 2015
Later among the works it cites.
Deep learning in neural networks: An overview
Schmidhuber, Jürgen · 2015
Later among the works it cites.
Trust region policy optimization
Schulman, John, Levine, Sergey, Moritz, Philipp, Jordan, Michael I, and Abbeel, Pieter · 2015
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, Karen and Zisserman, Andrew · 2015
Later among the works it cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, Bradly C, Levine, Sergey, and Abbeel, Pieter · 2015
Later among the works it cites.
Memory networks
Weston, Jason, Chopra, Sumit, and Bordes, Antoine · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2015
Cited alongside, same era.
Deep recurrent q-learning for partially observable mdps
Hausknecht, Matthew and Stone, Peter · 2015
Cited alongside, same era.
Inferring algorithmic patterns with stack-augmented recurrent nets
Joulin, Armand and Mikolov, Tomas · 2015
Cited alongside, same era.
Deep learning
LeCun, Yann, Bengio, Yoshua, and Hinton, Geoffrey · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Cited alongside, same era.
Language understanding for text-based games using deep reinforcement learning
Narasimhan, Karthik, Kulkarni, Tejas, and Barzilay, Regina · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, Junhyuk, Guo, Xiaoxiao, Lee, Honglak, Lewis, Richard L, and Singh, Satinder · 2015
Cited alongside, same era.
Reinforcement learning neural turing machines
Zaremba, Wojciech and Sutskever, Ilya · 2015
Later among the works it cites.
End-to-end training of deep visuomotor policies
Levine, Sergey, Finn, Chelsea, Darrell, Trevor, and Abbeel, Pieter · 2016
Closest in time.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P, Hunt, Jonathan J, Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adria Puigdomenech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy P, Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Closest in time.
Learning simple algorithms from examples
Zaremba, Wojciech, Mikolov, Tomas, Joulin, Armand, and Fergus, Rob · 2016
Closest in time.
Policy learning with continuous memory states for partially observed robotic control
Zhang, Marvin, Levine, Sergey, McCarthy, Zoe, Finn, Chelsea, and Abbeel, Pieter · 2016
Closest in time.