Fetching the paper…
Reading the bibliography…
In this paper, we study the sharpness of a deep learning (DL) loss landscape around local minima in order to reveal systematic mechanisms underlying the generalization abilities of DL models.
Some methods of speeding up the convergence of iteration methods
B.T. Polyak · 1964
Earlier work this paper cites.
Learning process in an asymmetric threshold network
Y. LeCun · 1986
Earlier work this paper cites.
Modèles connexionnistes de l’apprentissage
Y. LeCun · 1987
Earlier work this paper cites.
A theoretical framework for back-propagation
Y. LeCun, D. Touresky, G. Hinton, and T. Sejnowski · 1988
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Bayesian model comparison and backprop nets
D. J. C. MacKay · 1992
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Online algorithms and stochastic approximations
L. Bottou · 1998
Earlier work this paper cites.
Online algorithms and stochastic approximations
D. Saad · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian · 1999
Earlier work this paper cites.
Algorithmic stability and generalization performance
O. Bousquet and A. Elisseeff · 2001
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
Smooth minimization of non-smooth functions
Y. Nesterov · 2005
Earlier work this paper cites.
Decentralized resource allocation in dynamic networks of agents
H. Lakshmanan and D. Pucci De Farias · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Chapter 12. Significance and Measures of Association, 2011
R. E. Botsch · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Numerical continuation methods: an introduction , volume 13
E. L. Allgower and K. Georg · 2012
Earlier work this paper cites.
Randomized smoothing for stochastic optimization
J. C. Duchi, P. L. Bartlett, and M. J. Wainwright · 2012
Earlier work this paper cites.
Lecture 6.5: RmsProp: Divide the Gradient by a Running Average of Its Recent Magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
O. Bojar, C. Buck, C. Federmann, B. Haddow, P. Koehn, J. Leveling, C. Monz, P. Pecina, M. Post, H. Saint-Amand, R. Soricut, L. Specia, and A. Tamchyna · 2014
Earlier work this paper cites.
Distributed optimization of deeply nested systems
M. Carreira-Perpiñán and W. Wang · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Non-convex optimization, workshop at nips, 2015
A. Anandkumar, K. Chaudhuri, P. Liang, S. Oh, and U. N. Niranjan · 2015
Earlier work this paper cites.
Subdominant dense clusters allow for simple learning and high computational performance in neural networks with discrete synapses
C. Baldassi, A. Ingrosso, C. Lucibello, L. Saglietti, and R. Zecchina · 2015
Earlier work this paper cites.
On the energy landscape of deep networks
P. Chaudhari and S. Soatto · 2015
Earlier work this paper cites.
Natural neural networks
G. Desjardins, K. Simonyan, R. Pascanu, and K. Kavukcuoglu · 2015
Earlier work this paper cites.
Escaping from saddle points — online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
I. J. Goodfellow and O. Vinyals · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
B. Haeffele and R. Vidal · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Beating the Perils of Non-Convexity: Guaranteed Training of Neural Networks using Tensor Methods
M. Janzamin, H. Sedghi, and A. Anandkumar · 2015
Earlier work this paper cites.
Difference target propagation
D.-H. Lee, S. Zhang, A. Fischer, and Y. Bengio · 2015
Earlier work this paper cites.
On the link between Gaussian homotopy continuation and convex envelopes
H. Mobahi and J. Fisher III · 2015
Earlier work this paper cites.
Path-sgd: Path-normalized optimization in deep neural networks
B. Neyshabur, R. R. Salakhutdinov, and N. Srebro · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
Efficient approaches for escaping higher order saddle points in non-convex optimization
A. Anandkumar and R. Ge · 2016
Cited alongside, same era.
T. Cooijmans, N. Ballas, C. Laurent, and A. Courville · 2016
Cited alongside, same era.
C. Gulcehre, M. Moczulski, F. Visin, and Y. Bengio · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
K. Zhong, Z. Song, P. Jain, P. L. Bartlett, and I. S. Dhillon · 2017
Later among the works it cites.
A. Askari, G. Negiar, R. Sambharya, and L. El Ghaoui · 2018
Later among the works it cites.
Assessing the scalability of biologically-motivated deep learning algorithms and architectures
S. Bartunov, A. Santoro, B. Richards, L. Marris, G. E. Hinton, and T. Lillicrap · 2018
Later among the works it cites.
SGD learns over-parameterized networks that provably generalize on linearly separable data
A. Brutzkus, A. Globerson, E. Malach, and S. Shalev-Shwartz · 2018
Later among the works it cites.
Essentially no barriers in neural network energy landscape
F. Draxler, K. Veschgini, M. Salmhofer, and F. Hamprecht · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Deep learning without poor local minima
K. Kawaguchi · 2016
Cited alongside, same era.
Training recurrent neural networks by diffusion
H. Mobahi · 2016
Cited alongside, same era.
Deepfool: a simple and accurate method to fool deep neural networks
S.-M. Moosavi-Dezfooli, A. Fawzi, and P. Frossard · 2016
Cited alongside, same era.
On the quality of the initial basin in overspecified neural networks
I. Safran and O. Shamir · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
T. Salimans and D. Kingma · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
D. Soudry and Y. Carmon · 2016
Cited alongside, same era.
On the power of over-parametrization in neural networks with quadratic activation
S. Du and J. Lee · 2018
Later among the works it cites.
Decoupling backpropagation using constrained optimization methods
A. Gotmare, T. Valentin, J. Brea, and M. Jaggi · 2018
Later among the works it cites.
Deepcloak: Adversarial crafting as a defensive measure to cloak processes
M. S. Inci, T. Eisenbarth, and B. Sunar · 2018
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson · 2018
Later among the works it cites.
Finding flatter minima with sgd
S. Jastrzebski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2018
Later among the works it cites.
A proximal block coordinate descent algorithm for deep neural network training
T. T.-K. Lau, J. Zeng, B. Wu, and Y. Yao · 2018
Later among the works it cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Y. Li, T. Ma, and H. Zhang · 2018
Later among the works it cites.
On the stability and convergence of stochastic gradient descent with momentum
A. Ramezani-Kebrya, A. Khisti, and B. Liang · 2018
Later among the works it cites.
Empirical analysis of the hessian of over-parametrized neural networks
L. Sagun, U. Evci, V. Ugur Güney, Y. Dauphin, and L. Bottou · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
N. Shazeer and M. Stern · 2018
Later among the works it cites.
Smoothout: Smoothing out sharp minima for generalization in large-batch deep learning
W. Wen, Y. Wang, F. Yan, C. Xu, Y. Chen, and H. Li · 2018
Later among the works it cites.
Global optimality conditions for deep neural networks
C. Yun, S. Sra, and A. Jadbabaie · 2018
Later among the works it cites.
Global convergence in deep learning with variable splitting via the kurdyka-lojasiewicz property
J. Zeng, T. T.-K. Lau, S. Lin, and Y. Yao · 2018
Later among the works it cites.
A convergence analysis of gradient descent for deep linear neural networks
S. Arora, N. Cohen, N. Golowich, and W. Hu · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
M. Belkin, D. Hsu, S. Ma, and S. Mandal · 2019
Later among the works it cites.
Beyond backprop: Online alternating minimization with auxiliary variables
A. Choromanska, B. Cowen, S. Kumaravel, R. Luss, M. Rigotti, I. Rish, P. Diachille, V. Gurev, B. Kingsbury, R. Tejwani, and D. Bouneffouf · 2019
Later among the works it cites.
Lsalsa: accelerated source separation via learned sparse coding
B. Cowen, A. Nandini Saridena, and A. Choromanska · 2019
Later among the works it cites.
Autoaugment: Learning augmentation strategies from data
E. D. Cubuk, B. Zoph, D. Mane, V. Vasudevan, and Q. V. Le · 2019
Later among the works it cites.
An investigation into neural net optimization via hessian eigenvalue density
B. Ghorbani, S. Krishnan, and Y. Xiao · 2019
Later among the works it cites.
Training deep networks with stochastic gradient normalized by layerwise adaptive second moments
B. Ginsburg, P. Castonguay, O. Hrinchuk, O. Kuchaiev, R. Leary, V. Lavrukhin, J. Li, H. Nguyen, Y. Zhang, and J. M. Cohen · 2019
Later among the works it cites.
Gradient noise convolution (gnc): Smoothing loss function for distributed large-batch sgd
K. Haruki, T. Suzuki, Y. Hamakawa, T. Toda, R. Sakai, M. Ozawa, and M. Kimura · 2019
Later among the works it cites.
Fisher-rao metric, geometry, and complexity of neural networks
T. Liang, T. Poggio, A. Rakhlin, and J. Stokes · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
L. Luo, Y. Xiong, Y. Liu, and X. Sun · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
U. Simsekli, L. Sagun, and M. Gürbüzbalaban · 2019
Later among the works it cites.
Stability of stochastic gradient descent on nonsmooth convex losses
R. Bassily, V. Feldman, C. Guzmán, and K. Talwar · 2020
Later among the works it cites.
In search of robust measures of generalization
G. K. Dziugaite, A. Drouin, B. Neal, N. Rajkumar, E. Caballero, L. Wang, I. Mitliagkas, and D. M Roy · 2020
Later among the works it cites.
How neural networks find generalizable solutions: Self-tuned annealing in deep learning
Y. Feng and Y. Tu · 2020
Later among the works it cites.
Fantastic generalization measures and where to find them
Y. Jiang, B. Neyshabur, H. Mobahi, D. Krishnan, and S. Bengio · 2020
Later among the works it cites.
Extrapolation for large-batch training in deep learning
T. Lin, L. Kong, S. Stich, and M. Jaggi · 2020
Later among the works it cites.
Rethinking parameter counting in deep models: Effective dimensionality revisited
W. J. Maddox, G. Benton, and A. G. Wilson · 2020
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
P. Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, and I. Sutskever · 2020
Later among the works it cites.
Pure and spurious critical points: a geometric study of linear networks
M. Trager, K. Kohn, and J. Bruna · 2020
Later among the works it cites.
SWAD: Domain generalization by seeking flat minima
J. Cha, S. Chun, K. Lee, H.-C. Cho, S. Park, Y. Lee, and S. Park · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur · 2021
Later among the works it cites.
Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks
J. Kwon, J. Kim, H. Park, and I. K. Choi · 2021
Later among the works it cites.
Entropic gradient descent algorithms and wide flat minima
F. Pittorino, C. Lucibello, C. Feinauer, G. Perugini, C. Baldassi, E. Demyanenko, and R. Zecchina · 2021
Later among the works it cites.