Fetching the paper…
Reading the bibliography…
Training Deep Neural Networks is complicated by the fact that the distribution of each layer's inputs changes during training, as the parameters of the previous layers change.
Independent component analysis: Algorithms and applications
Hyvärinen, A. and Oja, E · 2000
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Shimodaira, Hidetoshi · 2000
Earlier work this paper cites.
A literature survey on domain adaptation of statistical classifiers, 2008
Jiang, Jing · 2008
Earlier work this paper cites.
Nonlinear image representation using divisive normalization
Lyu, S and Simoncelli, E P · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Bengio, Yoshua and Glorot, Xavier · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, Vinod and Hinton, Geoffrey E · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2011
Earlier work this paper cites.
A convergence analysis of log-linear training
Wiesler, Simon and Ney, Hermann · 2011
Cited alongside, same era.
Large scale distributed deep networks
Dean, Jeffrey, Corrado, Greg S., Monga, Rajat, Chen, Kai, Devin, Matthieu, Le, Quoc V., Mao, Mark Z., Ranzato, Marc’Aurelio, Senior, Andrew, Tucker, Paul, Yang, Ke, and Ng, Andrew Y · 2012
Cited alongside, same era.
Deep learning made easier by linear transformations in perceptrons
Raiko, Tapani, Valpola, Harri, and LeCun, Yann · 2012
Cited alongside, same era.
Knowledge matters: Importance of prior information for optimization
Gülçehre, Çaglar and Bengio, Yoshua · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2013
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Parallel training of deep neural networks with natural gradient and parameter averaging
Povey, Daniel, Zhang, Xiaohui, and Khudanpur, Sanjeev · 2014
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge, 2014
Russakovsky, Olga, Deng, Jia, Su, Hao, Krause, Jonathan, Satheesh, Sanjeev, Ma, Sean, Huang, Zhiheng, Karpathy, Andrej, Khosla, Aditya, Bernstein, Michael, Berg, Alexander C., and Fei-Fei, Li · 2014
Later among the works it cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, Nitish, Hinton, Geoffrey, Krizhevsky, Alex, Sutskever, Ilya, and Salakhutdinov, Ruslan · 2014
Later among the works it cites.
Going deeper with convolutions
Szegedy, Christian, Liu, Wei, Jia, Yangqing, Sermanet, Pierre, Reed, Scott, Anguelov, Dragomir, Erhan, Dumitru, Vanhoucke, Vincent, and Rabinovich, Andrew · 2014
Later among the works it cites.
Mean-normalized stochastic gradient for large-scale deep learning
Wiesler, Simon, Richard, Alexander, Schlüter, Ralf, and Ney, Hermann · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Saxe, Andrew M., McClelland, James L., and Ganguli, Surya · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Sutskever, Ilya, Martens, James, Dahl, George E., and Hinton, Geoffrey E · 2013
Cited alongside, same era.
Natural neural networks
Desjardins, Guillaume and Kavukcuoglu, Koray
Cited in the paper.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P
Cited in the paper.
Efficient backprop
LeCun, Y., Bottou, L., Orr, G., and Muller, K
Cited in the paper.
Later among the works it cites.
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Closest in time.
Deep image: Scaling up image recognition, 2015
Wu, Ren, Yan, Shengen, Shan, Yi, Dang, Qingqing, and Sun, Gang · 2015
Closest in time.