Fetching the paper…
Reading the bibliography…
It has been postulated that a good representation is one that disentangles the underlying explanatory factors of variation.
Likelilood ratio gradient estimation: an overview
Peter W Glynn · 1987
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
The Helmholtz machine
Peter Dayan, Geoffrey E Hinton, Radford M Neal, and Richard S Zemel · 1995
Earlier work this paper cites.
The MNIST database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Temporal abstraction in reinforcement learning
Doina Precup · 2000
Earlier work this paper cites.
Using MDP Characteristics to Guide Exploration in Reinforcement Learning
Bohdana Ratitch and Doina Precup · 2003
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Learning deep architectures for AI
Yoshua Bengio · 2009
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Modayil J. Delp M. Degris T. Pilarski P. M. White A. Precup-D. Sutton, R. S · 2011
Cited alongside, same era.
Reconstructing constructivism: Causal models, Bayesian learning mechanisms and the theory theory
Alison Gopnik and Henry M. Wellman · 2012
Cited alongside, same era.
Complex Valued Artificial Recurrent Neural Network as a Novel Approach to Model the Perceptual Binding Problem
Alexey Minin, Alois Knoll, Hans-Georg Zimmermann, AG Siemens, and LLC Siemens · 2012
Cited alongside, same era.
NICE: Non-linear Independent Components Estimation
Laurent Dinh, David Krueger, and Yoshua Bengio · 2014
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Later among the works it cites.
MazeBase: A sandbox for learning from games
Sainbayar Sukhbaatar, Arthur Szlam, Gabriel Synnaeve, Soumith Chintala, and Rob Fergus · 2015
Later among the works it cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2016
Later among the works it cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Later among the works it cites.
Tagger: Deep unsupervised perceptual grouping
Klaus Greff, Antti Rasmus, Mathias Berglund, Tele Hao, Harri Valpola, and Juergen Schmidhuber · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generative Adversarial Networks
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Auto-encoding variational Bayes
Durk P. Kingma and Max Welling · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Unsupervised Feature Extraction by Time-Contrastive Learning and Nonlinear ICA
Aapo Hyvarinen and Hiroshi Morioka · 2016
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Deep successor reinforcement learning
Tejas D Kulkarni, Ardavan Saeedi, Simanta Gautam, and Samuel J Gershman · 2016
Later among the works it cites.
The Predictron: End-To-End Learning and Planning
Matteo Hessel Tom Schaul Arthur Guez Tim Harley Gabriel Dulac-Arnold David Reichert Neil Rabinowitz Andre Barreto Thomas Degris David Silver, Hado van Hasselt · 2017
Closest in time.