Fetching the paper…
Reading the bibliography…
Recent work in deep reinforcement learning has allowed algorithms to learn complex tasks such as Atari 2600 games just from the reward provided by the game, but these algorithms presently require millions of training steps in order to learn, making them approximately five orders of magnitude slower than humans.
Activision, “Seaquest.” ROM cartridge, 1983
1983
Earlier work this paper cites.
A. L. Brown and M. J. Kane, “Preschool children can learn to transfer: Learning to learn and learning from example,” Cognitive Psychology
1988
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine learning
1992
Earlier work this paper cites.
D. H. Wolpert and W. G. Macready, “No free lunch theorems for optimization,” IEEE transactions on evolutionary computation
1997
Earlier work this paper cites.
J. Schmidhuber, “Discovering neural nets with low kolmogorov complexity and high generalization capability,” Neural Networks
1997
Earlier work this paper cites.
K. Ferguson and S. Mahadevan, “Proto-transfer learning in markov decision processes using spectral methods,” Computer Science Department Faculty Publication Series
2006
Earlier work this paper cites.
S. J. Pan, J. T. Kwok, and Q. Yang, “Transfer learning via dimensionality reduction.,” in AAAI
2008
Cited alongside, same era.
L. Torrey and J. Shavlik, “Transfer learning,” Handbook of Research on Machine Learning Applications and Trends: Algorithms, Methods, and Techniques
2009
Cited alongside, same era.
M. E. Taylor and P. Stone, “Transfer learning for reinforcement learning domains: A survey,” Journal of Machine Learning Research
2009
Cited alongside, same era.
S. J. Pan and Q. Yang, “A survey on transfer learning,” IEEE Transactions on knowledge and data engineering
2010
Cited alongside, same era.
2013
Cited alongside, same era.
Y. Du, V. Gabriel, J. Irwin, and M. E. Taylor, “Initial progress in transfer for deep reinforcement learning algorithms,” in Proceedings of Deep Reinforcement Learning: Frontiers and Challenges Workshop, New York City, NY, USA
2016
Later among the works it cites.
A. Tamar, Y. Wu, G. Thomas, S. Levine, and P. Abbeel, “Value iteration networks,” in Advances in Neural Information Processing Systems
2016
Later among the works it cites.
2016
Later among the works it cites.
M. Plappert, “keras-rl.” https://github.com/matthiasplappert/keras-rl, 2016
2016
Later among the works it cites.
B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman, “Building machines that learn and think like people,” Behavioral and Brain Sciences
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning.,” in AAAI
2016
Cited alongside, same era.
2016
Later among the works it cites.