Fetching the paper…
Reading the bibliography…
While the authors of Batch Normalization (BN) identify and address an important problem involved in training deep networks-- \textit{Internal Covariate Shift}-- the current solution has certain drawbacks.
Receptive fields of single neurones in the cat’s striate cortex
Hubel, D. H. and Wiesel, T. N · 1959
Earlier work this paper cites.
Distributed representations
Hinton, Geoffrey E · 1984
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, David E., Hinton, Geoffrey E., and Williams, Ronald J · 1986
Earlier work this paper cites.
Auto-association by multilayer perceptrons and singular value decomposition
Bourlard, H. and Kamp, Y · 1988
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: a strategy employed by v1
Olshausen, Bruno A. and Fieldt, David J · 1997
Earlier work this paper cites.
Where do you know what you know? the representation of semantic knowledge in the human brain
Patterson, Karalyn, Nestor, Peter, and Rogers, Timothy · 2007
Earlier work this paper cites.
Fast inference in sparse coding algorithms with applications to object recognition
Kavukcuoglu, Koray and Lecun, Yann · 2008
Earlier work this paper cites.
Sparse deep belief net model for visual area v2
Lee, Honglak, Ekanadham, Chaitanya, and Ng, Andrew Y · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Vincent, Pascal, Larochelle, Hugo, Bengio, Yoshua, and Manzagol, Pierre-Antoine · 2008
Cited alongside, same era.
Learning Multiple Layers of Features from Tiny Images
Krizhevsky, Alex · 2009
Cited alongside, same era.
Robust face recognition via sparse representation
Wright, J., Yang, A.Y., Ganesh, A., Sastry, S.S., and Ma, Yi · 2009
Cited alongside, same era.
Linear spatial pyramid matching using sparse coding for image classification
Yang, Jianchao, Yu, Kai, Gong, Yihong, and Huang, Thomas · 2009
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Glorot, Xavier and Bengio, Yoshua · 2010
Cited alongside, same era.
Rectified linear units improve restricted boltzmann machines
Nair, Vinod and Hinton, Geoffrey E · 2010
Sparse autoencoder
Ng, Andrew · 2011
Later among the works it cites.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, Geoffrey, Srivastava, Nitish, Krizhevsky, Alex, Sutskever, Ilya, and Salakhutdinov, Ruslan · 2012
Later among the works it cites.
Better mixing via deep representations
Bengio, Yoshua, Mesnil, Grégoire, Dauphin, Yann, and Rifai, Salah · 2013
Later among the works it cites.
Maxout networks
Goodfellow, Ian J., Warde-Farley, David, Mirza, Mehdi, Courville, Aaron C., and Bengio, Yoshua · 2013
Later among the works it cites.
Why does the unsupervised pretraining encourage moderate-sparseness?
Li, Jun, Luo, Wei, Yang, Jian, and Yuan, Xiaotong · 2013
Later among the works it cites.
Marginalized denoising auto-encoders for nonlinear representations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep sparse rectifier neural networks
Glorot, Xavier, Bordes, Antoine, and Bengio, Yoshua · 2011
Cited alongside, same era.
The MNIST database of handwritten digits
Lecun, Yann and Cortes, Corinna
Cited in the paper.
Higher order contractive auto-encoder
Rifai, Salah, Mesnil, Grégoire, Vincent, Pascal, Muller, Xavier, Bengio, Yoshua, Dauphin, Yann, and Glorot, Xavier
Cited in the paper.
Contractive auto-encoders: Explicit invariance during feature extraction
Rifai, Salah, Vincent, Pascal, Muller, Xavier, Glorot, Xavier, and Bengio, Yoshua
Cited in the paper.
Chen, Minmin, Weinberger, Kilian Q., Sha, Fei, and Bengio, Yoshua · 2014
Later among the works it cites.
Zero-bias autoencoders and the benefits of co-adapting features
Memisevic, Roland, Konda, Kishore Reddy, and Krueger, David · 2014
Later among the works it cites.