Fetching the paper…
Reading the bibliography…
Finding features that disentangle the different causes of variation in real data is a difficult task, that has nonetheless received considerable attention in static domains like natural images.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, Peter · 1993
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, Richard S, Precup, Doina, and Singh, Satinder · 1999
Earlier work this paper cites.
Temporal abstraction in reinforcement learning
Precup, Doina · 2000
Earlier work this paper cites.
Using mdp characteristics to guide exploration in reinforcement learning
Ratitch, B. and Precup, D · 2003
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, Geoffrey E and Salakhutdinov, Ruslan R · 2006
Cited alongside, same era.
Learning deep architectures for AI
Bengio, Yoshua · 2009
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil J. Delp M. Degris T. Pilarski P. M. White A. Precup-D · 2011
Cited alongside, same era.
Complex valued artificial recurrent neural network as a novel approach to model the perceptual binding problem
Minin, Alexey, Knoll, Alois, Zimmermann, Hans-Georg, Siemens, AG, and Siemens, LLC · 2012
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, Junhyuk, Guo, Xiaoxiao, Lee, Honglak, Lewis, Richard L, and Singh, Satinder · 2015
Cited alongside, same era.
Reconstructing constructivism: Causal models, bayesian learning mechanisms and the theory theory
Gopnik, A. and Wellman, H. M
Cited in the paper.
The option-critic architecture
Bacon, Pierre-Luc, Harb, Jean, and Precup, Doina · 2016
Later among the works it cites.
Tagger: Deep unsupervised perceptual grouping
Greff, Klaus, Rasmus, Antti, Berglund, Mathias, Hao, Tele, Valpola, Harri, and Schmidhuber, Juergen · 2016
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, Max, Mnih, Volodymyr, Czarnecki, Wojciech Marian, Schaul, Tom, Leibo, Joel Z, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Deep successor reinforcement learning
Kulkarni, Tejas D, Saeedi, Ardavan, Gautam, Simanta, and Gershman, Samuel J · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Silver, Hado van Hasselt, Matteo Hessel Tom Schaul Arthur Guez Tim Harley Gabriel Dulac-Arnold David Reichert Neil Rabinowitz Andre Barreto Thomas Degris · 2017
Closest in time.