Fetching the paper…
Reading the bibliography…
The Gauss-Newton (GN) matrix plays an important role in machine learning, most evident in its use as a preconditioning matrix for a wide family of popular adaptive methods to speed up optimization.
Das asymptotische verteilungsgesetz der eigenwerte linearer partieller differentialgleichungen (mit einer anwendung auf die theorie der hohlraumstrahlung)
Hermann Weyl · 1912
Earlier work this paper cites.
The smallest eigenvalue of a large dimensional wishart matrix
Jack W Silverstein · 1985
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla · 1989
Earlier work this paper cites.
Principles of risk minimization for learning theory
Vladimir Vapnik · 1991
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
Babak Hassibi, David G Stork, and Gregory J Wolff · 1993
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N. Schraudolph · 2002
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Earlier work this paper cites.
Layer normalization, 2016
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Shakeout: A new regularized deep neural network training scheme
Guoliang Kang, Jun Li, and Dacheng Tao · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Earlier work this paper cites.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data, 2017
Gintare Karolina Dziugaite and Daniel M. Roy · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
Jeffrey Pennington and Yasaman Bahri · 2017
Cited alongside, same era.
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Cited alongside, same era.
Group normalization
Yuxin Wu and Kaiming He · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Cited alongside, same era.
Hessian-based analysis of large batch training and robustness to adversaries
Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer, and Michael W Mahoney · 2018
Woodfisher: Efficient second-order approximation for neural network compression
Sidak Pal Singh and Dan Alistarh · 2020
Later among the works it cites.
Understanding black-box predictions via influence functions, 2020
Pang Wei Koh and Percy Liang · 2020
Later among the works it cites.
Dissecting hessian: Understanding common structure of hessian in neural networks
Yikai Wu, Xingyu Zhu, Chenwei Wu, Annie Wang, and Rong Ge · 2020
Later among the works it cites.
Provable benefit of orthogonal initialization in optimizing deep linear networks
Wei Hu, Lechao Xiao, and Jeffrey Pennington · 2020
Later among the works it cites.
Tensor programs ii: Neural tangent kernel for any architecture
Greg Yang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Cited alongside, same era.
The difficulty of training sparse neural networks
Utku Evci, Fabian Pedregosa, Aidan Gomez, and Erich Elsen · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2019
Cited alongside, same era.
An investigation into neural net optimization via hessian eigenvalue density, 2019
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Cited alongside, same era.
Analytic insights into structure and rank of neural network hessian maps
Sidak Pal Singh, Gregor Bachmann, and Thomas Hofmann · 2021
Later among the works it cites.
Flatness is a false friend, 2021
Diego Granziol · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization, 2021
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy M Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Later among the works it cites.
Adahessian: An adaptive second order optimizer for machine learning, 2021
Zhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa, Kurt Keutzer, and Michael W. Mahoney · 2021
Later among the works it cites.
Hessian eigenspectra of more realistic nonlinear models
Zhenyu Liao and Michael W Mahoney · 2021
Later among the works it cites.
Feature learning and signal propagation in deep neural networks
Yizhang Lou, Chris E Mingard, and Soufiane Hayou · 2022
Later among the works it cites.
Phenomenology of double descent in finite-width neural networks
Sidak Pal Singh, Aurelien Lucchi, Thomas Hofmann, and Bernhard Schölkopf · 2022
Later among the works it cites.
When and why pinns fail to train: A neural tangent kernel perspective
Sifan Wang, Xinling Yu, and Paris Perdikaris · 2022
Later among the works it cites.
Tight bounds on the smallest eigenvalue of the neural tangent kernel for deep relu networks, 2022
Quynh Nguyen, Marco Mondelli, and Guido Montufar · 2022
Later among the works it cites.
Sophia: A scalable stochastic second-order optimizer for language model pre-training, 2023
Hong Liu, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma · 2023
Later among the works it cites.
A modern look at the relationship between sharpness and generalization
Maksym Andriushchenko, Francesco Croce, Maximilian Müller, Matthias Hein, and Nicolas Flammarion · 2023
Later among the works it cites.
Wu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov, Agustinus Kristiadi, Richard E Turner, and Alireza Makhzani · 2023
Later among the works it cites.
Optimal brain compression: A framework for accurate post-training quantization and pruning, 2023
Elias Frantar, Sidak Pal Singh, and Dan Alistarh · 2023
Later among the works it cites.
Studying large language model generalization with influence functions, 2023
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, Evan Hubinger, Kamilė Lukošiūtė, Karina Nguyen, Nicholas Joseph, Sam McCandlish, Jared Kaplan, and Samuel R. Bowman · 2023
Later among the works it cites.
The hessian perspective into the nature of convolutional neural networks
Sidak Pal Singh, Thomas Hofmann, and Bernhard Schölkopf · 2023
Later among the works it cites.
Relu soothes the ntk condition number and accelerates optimization for wide neural networks
Chaoyue Liu and Like Hui · 2023
Later among the works it cites.
Average gradient outer product as a mechanism for deep neural collapse
Daniel Beaglehole, Peter Súkeník, Marco Mondelli, and Mikhail Belkin · 2024
Closest in time.