Fetching the paper…
Reading the bibliography…
When the parameters are independently and identically distributed (initialized) neural networks exhibit undesirable properties that emerge as the number of layers increases, e.g.
Traditional and heavy-tailed self regularization in neural network models
Martin, C. H. and Mahoney, M. W. (2019) · 1901
Earlier work this paper cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Simsekli, U., Sagun, L., and Gurbuzbalaban, M. (2019) · 1901
Earlier work this paper cites.
Training dynamics of deep networks using stochastic gradient descent via neural tangent kernel
Hayou, S., Doucet, A., and Rousseau, J. (2019b) · 1905
Earlier work this paper cites.
Arch models as diffusion approximations
Nelson, D. B. (1990) · 1990
Earlier work this paper cites.
Arch models as diffusion approximations
Nelson, D. B. (1990) · 1990
Earlier work this paper cites.
Numerical Solution of Stochastic Differential Equations
Kloeden, P. E. and Platen, E. (1992) · 1992
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Neal, R. M. (1995) · 1995
Earlier work this paper cites.
The mnist database of handwritten digits
LeCun, Y. (1998) · 1998
Earlier work this paper cites.
Convergence of Probability Measures
Billingsley, P. (1999) · 1999
Earlier work this paper cites.
Matrix variate distributions
Gupta, A. K. and Nagar, D. K. (1999) · 1999
Earlier work this paper cites.
Continuous Martingales and Brownian Motion
Revuz, D. and Yor, M. (1999) · 1999
Earlier work this paper cites.
Stochastic Differential Equations: An Introduction with Applications
Øksendal, B. (2003) · 2003
Earlier work this paper cites.
Stochastic Differential Equations: An Introduction with Applications
Øksendal, B. (2003) · 2003
Earlier work this paper cites.
Multidimensional diffusion processes
Stroock, D. W. and Varadhan, S. S. (2006) · 2006
Cited alongside, same era.
Multidimensional diffusion processes
Stroock, D. W. and Varadhan, S. S. (2006) · 2006
Cited alongside, same era.
Stochastic calculus for fractional Brownian motion and applications
Biagini, F., Hu, Y., Øksendal, B., and Zhang, T. (2008) · 2008
Cited alongside, same era.
Markov processes: characterization and convergence
Ethier, S. N. and Kurtz, T. G. (2009) · 2009
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y. (2010) · 2010
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C. (2018) · 2018
Later among the works it cites.
Deep neural networks as gaussian processes
Lee, J., Sohl-dickstein, J., Pennington, J., Novak, R., Schoenholz, S., and Bahri, Y. (2018) · 2018
Later among the works it cites.
Gaussian process behaviour in wide deep neural networks
Matthews, A. G. d. G., Rowland, M., Hron, J., Turner, R. E., and Ghahramani, Z. (2018) · 2018
Later among the works it cites.
Neural ordinary differential equations
Chen, T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K. (2018) · 2018
Later among the works it cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Hanin, B. (2018) · 2018
Later among the works it cites.
How to start training: The effect of initialization and architecture
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G. (2015) · 2015
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Poole, B., Lahiri, S., Raghu, M., Sohl-Dickstein, J., and Ganguli, S. (2016) · 2016
Cited alongside, same era.
Searching for activation functions
Ramachandran, P., Zoph, B., and Le, Q. V. (2017) · 2017
Cited alongside, same era.
Deep information propagation
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J. (2017) · 2017
Cited alongside, same era.
Mean field residual networks: On the edge of chaos
Yang, G. and Schoenholz, S. (2017) · 2017
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Pennington, J., Schoenholz, S., and Ganguli, S. (2017) · 2017
Cited alongside, same era.
Hanin, B. and Rolnick, D. (2018) · 2018
Later among the works it cites.
The emergence of spectral universality in deep networks
Pennington, J., Schoenholz, S. S., and Ganguli, S. (2018) · 2018
Later among the works it cites.
Handbook of approximate Bayesian computation
Sisson, S. A., Fan, Y., and Beaumont, M. (2018) · 2018
Later among the works it cites.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R. (2019) · 2019
Closest in time.
Deep convolutional networks as shallow gaussian processes
Garriga-Alonso, A., Rasmussen, C. E., and Aitchison, L. (2019) · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S. S., Bahri, Y., Sohl-Dickstein, J., and Pennington, J. (2019) · 2019
Closest in time.
Residual learning without normalization via better initialization
Zhang, H., Dauphin, Y. N., and Ma, T. (2019) · 2019
Closest in time.
Initialization of relus for dynamical isometry
Burkholz, R. and Dubatovka, A. (2019) · 2019
Closest in time.