Fetching the paper…
Reading the bibliography…
The early layers of a deep neural net have the fewest parameters, but take up the most computation.
Greedy layer-wise training of deep networks
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle · 2007
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P.-A. Manzagol · 2010
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
N. Srivastava, G.E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Deep networks with stochastic depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger · 2016
Cited alongside, same era.
Wide residual networks
S. Zagoruyko and N. Komodakis · 2016
Later among the works it cites.
Shake-shake regularization of 3-branch residual networks
X. Gastaldi · 2017
Closest in time.
Densely connected convolutional networks
G. Huang, Z. Liu, K.Q. Weinberger, and L. van der Maaten · 2017
Closest in time.
Sgdr: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…