Fetching the paper…
Reading the bibliography…
The choice of initial learning rate can have a profound effect on the performance of deep networks.
The effect of network width on stochastic gradient descent and generalization: an empirical study
Park, D. S., Sohl-Dickstein, J., Le, Q. V., and Smith, S. L · 1905
Earlier work this paper cites.
Simple mathematical models with very complicated dynamics
May, R. M · 1976
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Sgd learns the conjugate kernel class of the network
Daniely, A · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Earlier work this paper cites.
Stochastic gradient descent as approximate bayesian inference
Mandt, S., Hoffman, M. D., and Blei, D. M · 2017
Earlier work this paper cites.
Don’t Decay the Learning Rate, Increase the Batch Size
Smith, S. L., Kindermans, P.-J., Ying, C., and Le, Q. V · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., and Wanderman-Milne, S · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Deep neural networks as gaussian processes
Lee, J., Bahri, Y., Novak, R., Schoenholz, S., Pennington, J., and Sohl-dickstein, J · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P.-M · 2018
Cited alongside, same era.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
Rotskoff, G. and Vanden-Eijnden, E · 2018
Cited alongside, same era.
Mean field analysis of neural networks
Sirignano, J. and Spiliopoulos, K · 2018
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Smith, S. L. and Le, Q. V · 2018
Cited alongside, same era.
Stochastic natural gradient descent draws posterior samples in function space
Towards explaining the regularization effect of initial large learning rate in training neural networks
Li, Y., Wei, C., and Ma, T · 2019
Later among the works it cites.
Bayesian deep convolutional networks with many channels are gaussian processes
Novak, R., Xiao, L., Bahri, Y., Lee, J., Yang, G., Abolafia, D. A., Pennington, J., and Sohl-dickstein, J · 2019
Later among the works it cites.
Kernel and deep regimes in overparametrized models
Woodworth, B., Gunasekar, S., Lee, J., Soudry, D., and Srebro, N · 2019
Later among the works it cites.
Disentangling trainability and generalization in deep learning, 2019
Xiao, L., Pennington, J., and Schoenholz, S. S · 2019
Later among the works it cites.
Asymptotics of wide networks from feynman diagrams
Dyer, E. and Gur-Ari, G · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Smith, S. L., Duckworth, D., Rezchikov, S., Le, Q. V., and Sohl-Dickstein, J · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R. R., and Wang, R · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S. S., Lee, J. D., Li, H., Wang, L., and Zhai, X · 2019
Cited alongside, same era.
Dynamics of Deep Neural Networks and Neural Tangent Hierarchy
Huang, J. and Yau, H.-T · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J · 2019
Cited alongside, same era.
Frankle, J., Schwab, D. J., and Morcos, A. S · 2020
Closest in time.
The break-even point on optimization trajectories of deep neural networks
Jastrzebski, S., Szymczak, M., Fort, S., Arpit, D., Tabor, J., Cho, K., and Geras, K · 2020
Closest in time.
Fantastic generalization measures and where to find them
Jiang, Y., Neyshabur, B., Krishnan, D., Mobahi, H., and Bengio, S · 2020
Closest in time.
The two regimes of deep network training, 2020
Leclerc, G. and Madry, A · 2020
Closest in time.
Neural tangents: Fast and easy infinite neural networks in python
Novak, R., Xiao, L., Hron, J., Lee, J., Alemi, A. A., Sohl-Dickstein, J., and Schoenholz, S. S · 2020
Closest in time.
Xie, Z., Sato, I., and Sugiyama, M · 2020
Closest in time.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2020
Closest in time.