Fetching the paper…
Reading the bibliography…
We consider deep classifying neural networks.
Multinomial logistic regression algorithm
Böhning, D · 1992
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Visualizing data using t-sne
Maaten, L. v. d. and Hinton, G · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Mnist handwritten digit database
LeCun, Y., Cortes, C., and Burges, C · 2010
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Sagun, L., Bottou, L., and LeCun, Y · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Later among the works it cites.
The jamming transition as a paradigm to understand the loss landscape of deep neural networks
Geiger, M., Spigler, S., d’Ascoli, S., Sagun, L., Baity-Jesi, M., Biroli, G., and Wyart, M · 2018
Later among the works it cites.
Gradient descent happens in a tiny subspace
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Later among the works it cites.
On the relation between the sharpest directions of dnn loss and the sgd step length
Jastrzkebski, S., Kenton, Z., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A · 2018
Later among the works it cites.
The full spectrum of deep net hessians at scale: Dynamics with sample size
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
Pennington, J. and Bahri, Y · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2017
Cited alongside, same era.
Papyan, V · 2018
Later among the works it cites.
The spectrum of the fisher information matrix of a single-hidden-layer neural network
Pennington, J. and Worah, P · 2018
Later among the works it cites.
A jamming transition from under-to over-parametrization affects loss landscape and generalization
Spigler, S., Geiger, M., d’Ascoli, S., Sagun, L., Biroli, G., and Wyart, M · 2018
Later among the works it cites.
Fluctuation-dissipation relations for stochastic gradient descent
Yaida, S · 2018
Later among the works it cites.