Fetching the paper…
Reading the bibliography…
For a long time, designing neural architectures that exhibit high performance was considered a dark art that required expert hand-tuning.
Untersuchungen zu dynamischen neuronalen netzen
Hochreiter, S · 1991
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and P, F · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Shallow vs. deep sum-product networks
Delalleau, O. and Bengio, Y · 2011
Earlier work this paper cites.
On the representational efficiency of restricted boltzmann machines
Martens, J., Chattopadhya, A., Pitassi, T., and Zemel, R · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Earlier work this paper cites.
On the complexity of shallow and deep neural network classifiers
Bianchini, M. and Scarselli, F · 2014
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Montafur, G. F., Pascanu, R., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
The power of depth for feedforward neural networks
Shamir, O. and Eldan, R · 2015
Earlier work this paper cites.
Representation benefits of deep feedforward networks
Telgarsky, M · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
Arjovsky, M., Shah, A., and Bengio, Y · 2016
Cited alongside, same era.
Arpit, D., Zhou, Y., Kota, B. U., and Govindaraju, V · 2016
Cited alongside, same era.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Pennington, J., Schoenholz, S., and Ganguli, S · 2017
Later among the works it cites.
On the expressive power of deep neural networks
Raghu, M., Poole, B., Ganguli, S., Kleinberg, J., and Sohl-Dickstein, J · 2017
Later among the works it cites.
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J · 2017
Later among the works it cites.
Mean field residual networks: On the edge of chaos
Yang, G. and Schoenholz, S. S · 2017
Later among the works it cites.
Dynamical isometry and a mean field theory of rnns: Gating enables signal propagation in recurrent neural networks
Chen, M., Pennington, J., and Schoenholz, S. S · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Deep vs. shallow networks: An approximation theory perspective
Mhaskar, H. N. and Shamir, O · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Poole, B., Lahiri, S., Raghu, M., Sohl-Dickstein, J., and Ganguli, S · 2016
Cited alongside, same era.
The shattered gradients problem: If resnets are the answer, then what is the question?
Balduzzi, D., Frean, M., Leary, L., Lewis, J., Ma, K. W.-D., and McWilliams, B · 2017
Cited alongside, same era.
Parseval networks: Improving robustness to adversarial examples
Cisse, M., Bojanowski, P., Grave, E., Douphin, Y., and Usunier, N · 2017
Cited alongside, same era.
Deep pyramidal residual networks
Han, D., Kim, J., and Kim, J · 2017
Cited alongside, same era.
Self-normalizing neural networks
Klambauer, G., Unterthiner, T., Mayr, A., and Hochreiter, S · 2017
Cited alongside, same era.
Nonlinear random matrix theory for deep learning
Pennington, J. and Worah, P · 2017
Cited alongside, same era.
Closest in time.
Orthogonal recurrent neural networks with scaled cayley transform
Helfrich, K., WIllmott, D., and Ye, Q · 2018
Closest in time.
Sensitivity and generalization in neural networks: an empirical study
Novak, R., Bahri, Y., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J · 2018
Closest in time.
Philipp, G., Song, D., and Carbonell, J. G · 2018
Closest in time.
Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks
Xiao, L., Bahri, Y., Sohl-Dickstein, J., Schoenholz, S. S., and Pennington, J · 2018
Closest in time.
Deep mean field theory: Variance and width variation by layer as methods to control gradient explosion
Yang, G. and Schoenholz, S. S · 2018
Closest in time.
Predicting the generalization gap in deep networks with margin distributions
Jiang, Y., Krishnan, D., Mobahi, H., and Bengio, S · 2019
Closest in time.
A mean field theory of batch normalization
Yang, G., Pennington, J., Rao, V., Sohl-Dickstein, J., and Schoenholz, S. S · 2019
Closest in time.