Fetching the paper…
Reading the bibliography…
Analyzing geometric properties of high-dimensional loss functions, such as local curvature and the existence of other optima around a certain point in loss space, can help provide a better understanding of the interplay between neural network structure, implementation attributes, and learning performance.
E. P. Wigner, On the statistical distribution of the widths and spacings of nuclear resonance levels, in Mathematical Proceedings of the Cambridge Philosophical Society , Vol. 47 (Cambridge University Press, 1951) pp. 790–798
1951
Earlier work this paper cites.
V. A. Marčenko and L. A. Pastur, Distribution of eigenvalues for some sets of random matrices, Mathematics of the USSR-Sbornik 1
1967
Earlier work this paper cites.
R. B. Davies, Algorithm AS 155: The distribution of a linear combination of χ 2 \chi^{2} random variables, Applied Statistics , 323 (1980)
1980
Earlier work this paper cites.
K. Hornik, M. Stinchcombe, and H. White, Multilayer feedforward networks are universal approximators, Neural Networks 2
1989
Earlier work this paper cites.
P. Baldi and K. Hornik, Neural networks and principal component analysis: Learning from examples without local minima, Neural Networks 2
1989
Earlier work this paper cites.
M. F. Hutchinson, A stochastic estimator of the trace of the influence matrix for Laplacian smoothing splines, Communications in Statistics-Simulation and Computation 18
1989
Earlier work this paper cites.
K. Hornik, Approximation capabilities of multilayer feedforward networks, Neural Networks 4
1991
Earlier work this paper cites.
K.-C. Li, On principal hessian directions for data visualization and dimension reduction: Another application of Stein’s lemma, Journal of the American Statistical Association 87
1992
Earlier work this paper cites.
R. Hecht-Nielsen, Context vectors: general purpose approximate meaning representations self-organized from raw data, Computational Intelligence: Imitating Life, IEEE Press , 43 (1994)
1994
Earlier work this paper cites.
B. A. Pearlmutter, Fast exact multiplication by the Hessian, Neural Computation 6
1994
Earlier work this paper cites.
Z. Bai, G. Fahey, and G. Golub, Some large-scale matrix computation problems, Journal of Computational and Applied Mathematics 74
1996
Earlier work this paper cites.
S.-I. Amari, Natural gradient works efficiently in learning, Neural Computation 10
1998
Earlier work this paper cites.
R. B. Lehoucq, D. C. Sorensen, and C. Yang, ARPACK users’ guide: solution of large-scale eigenvalue problems with implicitly restarted Arnoldi methods (SIAM, 1998)
1998
Earlier work this paper cites.
S.-i. Amari, H. Park, and K. Fukumizu, Adaptive method of realizing natural gradient learning for multilayer perceptrons, Neural Computation 12
2000
Earlier work this paper cites.
H. Park, S.-I. Amari, and K. Fukumizu, Adaptive natural gradient learning algorithms for various stochastic models, Neural Networks 13
2000
Earlier work this paper cites.
B. Laurent and P. Massart, Adaptive estimation of a quadratic functional by model selection, Annals of Statistics , 1302 (2000)
2000
Earlier work this paper cites.
G. Casella and R. L. Berger, Statistical inference, 2nd edition (Duxbury Press, Pacific Grove, CA, 2002)
2002
Earlier work this paper cites.
N. N. Schraudolph, Fast curvature matrix-vector products for second-order gradient descent, Neural Computation 14
2002
Earlier work this paper cites.
J. M. Lee, Riemannian Manifolds: An Introduction to Curvature , Vol. 176 (Springer Science & Business Media, 2006)
2006
Earlier work this paper cites.
E. L. Lehmann and G. Casella, Theory of point estimation, 2nd edition (Springer-Verlag, New York, NY, USA, 2006)
2006
Earlier work this paper cites.
A. J. Bray and D. S. Dean, Statistics of critical points of Gaussian fields on large-dimensional spaces, Physical Review Letters 98
2007
Earlier work this paper cites.
J. Matoušek, On variants of the johnson-lindenstrauss lemma, Random Structures and Algorithms 33
2008
Earlier work this paper cites.
H. Avron and S. Toledo, Randomized algorithms for estimating the trace of an implicit symmetric positive semi-definite matrix, Journal of the ACM (JACM) 58
2011
Cited alongside, same era.
M. Berger and B. Gostiaux, Differential Geometry: Manifolds, Curves, and Surfaces , Vol. 115 (Springer Science & Business Media, 2012)
2012
Cited alongside, same era.
J. Bausch, On the efficient calculation of a linear combination of chi-square random variables with an application in counting string vacua, Journal of Physics A: Mathematical and Theoretical 46
2013
Cited alongside, same era.
Y. N. Dauphin, R. Pascanu, Ç. Gülçehre, K. Cho, S. Ganguli, and Y. Bengio, Identifying and attacking the saddle point problem in high-dimensional non-convex optimization, in Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada , edited by Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger (2014) pp. 2933–2941
D. Choi, C. J. Shallue, Z. Nado, J. Lee, C. J. Maddison, and G. E. Dahl, On empirical comparisons of optimizers for deep learning, International Conference on Learning Representations (2020)
2020
Later among the works it cites.
C. Baldassi, F. Pittorino, and R. Zecchina, Shaping the learning landscape in neural networks around wide flat minima, Proceedings of the National Academy of Sciences 117
2020
Later among the works it cites.
D. Wu, S.-T. Xia, and Y. Wang, Adversarial weight perturbation helps robust generalization, Advances in Neural Information Processing Systems 33
2020
Later among the works it cites.
Z. Yao, A. Gholami, K. Keutzer, and M. W. Mahoney, PyHessian: Neural networks through the lens of the Hessian, in 2020 IEEE International Conference on Big Data (Big data) (IEEE, 2020) pp. 581–590
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
I. J. Goodfellow and O. Vinyals, Qualitatively Characterizing Neural Network Optimization Problems, in 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings , edited by Y. Bengio and Y. LeCun (2015)
2015
Cited alongside, same era.
W. Kühnel, Differential Geometry , Vol. 77 (American Mathematical Society, 2015)
2015
Cited alongside, same era.
R. L. Burden, J. D. Faires, and A. M. Burden, Numerical Analysis; 10th edition (Cengage Learning, Boston, MA, 2015)
2015
Cited alongside, same era.
M. Hardt, B. Recht, and Y. Singer, Train faster, generalize better: Stability of stochastic gradient descent, in International Conference on Machine Learning (Proceedings of Machine Learning Research, 2016) pp. 1225–1234
2016
Cited alongside, same era.
S. Zagoruyko and N. Komodakis, Wide Residual Networks, in Proceedings of the British Machine Vision Conference 2016, BMVC 2016, York, UK, September 19-22, 2016 , edited by R. C. Wilson, E. R. Hancock, and W. A. P. Smith (BMVA Press, 2016)
2016
Cited alongside, same era.
A. C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht, The marginal value of adaptive gradient methods in machine learning, Advances in Neural Information Processing Systems 30
2017
Cited alongside, same era.
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang, On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima, in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings (2017)
2017
Cited alongside, same era.
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio, Sharp Minima Can Generalize For Deep Nets, in Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 , Proceedings of Machine Learning Research, Vol. 70, edited by D. Precup and Y. W. Teh (2017) pp. 1019–1028
2017
Cited alongside, same era.
2020
Later among the works it cites.
R. Karakida, S. Akaho, and S.-i. Amari, Universal statistics of Fisher information in deep neural networks: mean field approach, Journal of Statistical Mechanics: Theory and Experiment 2020
2020
Later among the works it cites.
Y. Cooper, Global minima of overparameterized neural networks, SIAM Journal on Mathematics of Data Science 3
2021
Later among the works it cites.
S. Park, C. Yun, J. Lee, and J. Shin, Minimum Width for Universal Approximation, in 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 (2021)
2021
Later among the works it cites.
F. Pittorino, C. Lucibello, C. Feinauer, G. Perugini, C. Baldassi, E. Demyanenko, and R. Zecchina, Entropic gradient descent algorithms and wide flat minima, Journal of Statistical Mechanics: Theory and Experiment 2021
2021
Later among the works it cites.
B. Adcock and N. Dexter, The gap between theory and practice in function approximation with deep neural networks, SIAM Journal on Mathematics of Data Science 3
2021
Later among the works it cites.
Z. Liao and M. W. Mahoney, Hessian eigenspectra of more realistic nonlinear models, Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, in 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 (OpenReview.net, 2021)
2021
Later among the works it cites.
S. Horoi, J. Huang, B. Rieck, G. Lajoie, G. Wolf, and S. Krishnaswamy, Exploring the Geometry and Topology of Neural Network Loss Landscapes, in Advances in Intelligent Data Analysis XX - 20th International Symposium on Intelligent Data Analysis, IDA 2022, Rennes, France, April 20-22, 2022, Proceedings , Lecture Notes in Computer Science, Vol. 13205, edited by T. Bouadi, É. Fromont, and E. Hüllermeier (Springer, 2022) pp. 171–184
2022
Closest in time.
L. Böttcher, GitLab repository, https://gitlab.com/ComputationalScience/loss-visualization (2022)
2022
Closest in time.
H. Phan, Pretrained TorchVision models on CIFAR-10 dataset, https://github.com/huyvnphan/PyTorch_CIFAR10 (2022)
2022
Closest in time.
H. Touvron, P. Bojanowski, M. Caron, M. Cord, A. El-Nouby, E. Grave, G. Izacard, A. Joulin, G. Synnaeve, J. Verbeek, et al. , ResMLP: Feedforward networks for image classification with data-efficient training, IEEE Transactions on Pattern Analysis and Machine Intelligence (2022)
2022
Closest in time.
T. Asikis, L. Böttcher, and N. Antulov-Fantulin, Neural ordinary differential equation control of dynamics on graphs, Physical Review Research 4
2022
Closest in time.
L. Böttcher, N. Antulov-Fantulin, and T. Asikis, AI Pontryagin or how artificial neural networks learn to control dynamical systems, Nature Communications 13
2022
Closest in time.
J. W. Baron, Eigenvalue spectra and stability of directed complex networks, Physical Review E 106
2022
Closest in time.
L. Böttcher, Video of loss evolution, https://vimeo.com/787174225 (2023)
2023
Closest in time.
L. Böttcher, Gradient-free training of neural ODEs for system identification and control using ensemble Kalman inversion, in ICML Workshop on New Frontiers in Learning, Control, and Dynamical Systems (2023)
2023
Closest in time.