Fetching the paper…
Reading the bibliography…
This paper studies how neural network architecture affects the speed of training.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Ensembles semi-analytiques
Lojasiewicz, S · 1965
Earlier work this paper cites.
Asymptotic Theory of Finite Dimensional Normed Spaces
Milman, V. D. and Schechtman, G · 1986
Earlier work this paper cites.
Towards faster stochastic gradient search
Darken, C. and Moody, J · 1992
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P · 1994
Earlier work this paper cites.
Convergence rate of incremental subgradient algorithms
Nedić, A. and Bertsekas, D · 2001
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Incremental gradient, subgradient, and proximal methods for convex optimization: A survey
Bertsekas, D. P · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E. and Bach, F. R · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Efficient backprop
LeCun, Y. A., Bottou, L., Orr, G. B., and Müller, K.-R · 2012
Earlier work this paper cites.
Topics in random matrix theory , volume 132
Tao, T · 2012
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
Boucheron, S., Lugosi, G., and Massart, P · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2013
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition
Schmidt, M. and Roux, N. L · 2013
Earlier work this paper cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Shamir, O. and Zhang, T · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Needell, D., Ward, R., and Srebro, N · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2016
Cited alongside, same era.
The power of depth for feedforward neural networks
Eldan, R. and Shamir, O · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Cited alongside, same era.
Second-order optimization for neural networks
Martens, J · 2016
Cited alongside, same era.
The loss landscape of overparameterized neural networks
Cooper, Y · 2018
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2018
Later among the works it cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Hanin, B · 2018
Later among the works it cites.
How to start training: The effect of initialization and architecture
Hanin, B. and Rolnick, D · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J · 2016
Cited alongside, same era.
Benefits of depth in neural networks
Telgarsky, M · 2016
Cited alongside, same era.
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
The shattered gradients problem: If resnets are the answer, then what is the question?
Balduzzi, D., Frean, M., Leary, L., Lewis, J., Ma, K. W.-D., and McWilliams, B · 2017
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2017
Cited alongside, same era.
Automated inference with adaptive batches
De, S., Yadav, A., Jacobs, D., and Goldstein, T · 2017
Cited alongside, same era.
Dziugaite, G. K. and Roy, D. M · 2017
Cited alongside, same era.
Nesterov, Y · 2018
Later among the works it cites.
Towards understanding the role of over-parametrization in generalization of neural networks
Neyshabur, B., Li, Z., Bhojanapalli, S., LeCun, Y., and Srebro, N · 2018
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Oymak, S. and Soltanolkotabi, M · 2018
Later among the works it cites.
The singular values of convolutional layers
Sedghi, H., Gupta, V., and Long, P. M · 2018
Later among the works it cites.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Vaswani, S., Bach, F., and Schmidt, M · 2018
Later among the works it cites.
High-dimensional probability: An introduction with applications in data science , volume 47
Vershynin, R · 2018
Later among the works it cites.
Xiao, L., Bahri, Y., Sohl-Dickstein, J., Schoenholz, S. S., and Pennington, J · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2018
Later among the works it cites.
Stiffness: A new perspective on generalization in neural networks
Fort, S., Nowak, P. K., and Narayanan, S · 2019
Closest in time.
An investigation into neural net optimization via hessian eigenvalue density
Ghorbani, B., Krishnan, S., and Xiao, Y · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S. S., Bahri, Y., Sohl-Dickstein, J., and Pennington, J · 2019
Closest in time.
Generalization in deep networks: The role of distance from initialization
Nagarajan, V. and Kolter, J. Z · 2019
Closest in time.
A mean field theory of batch normalization
Yang, G., Pennington, J., Rao, V., Sohl-Dickstein, J., and Schoenholz, S. S · 2019
Closest in time.
Batch normalization biases residual blocks towards the identity function in deep networks
De, S. and Smith, S. L · 2020
Closest in time.
On the generalization benefit of noise in stochastic gradient descent
Smith, S. L., Elsen, E., and De, S · 2020
Closest in time.