Fetching the paper…
Reading the bibliography…
Structured large matrices are prevalent in machine learning.
Discriminatory Analysis: Nonparametric Discrimination: Consistency Properties
Fix, E. and Hodges, J · 1951
Earlier work this paper cites.
Methods of conjugate gradients for solving linear systems
Hestenes, M. R. and Stiefel, E · 1952
Earlier work this paper cites.
Sur la division des corps matériels en parties
Steinhaus, H. et al · 1956
Earlier work this paper cites.
Applications of negative dimensional tensors
Penrose, R · 1971
Earlier work this paper cites.
Iterative methods for solving minimax problems
Evtushenko, Y. G · 1974
Earlier work this paper cites.
The influence curve and its role in robust estimation
Hampel, F. R · 1974
Earlier work this paper cites.
A fast ’monte-carlo cross-validation’ procedure for large least squares problems with noisy data
Girard, D. A · 1989
Earlier work this paper cites.
A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines
Hutchinson, M · 1989
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J., and Solla, S · 1989
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B. and Stork, D · 1992
Earlier work this paper cites.
A practical Bayesian framework for backpropagation networks
MacKay, D. J · 1992
Earlier work this paper cites.
Fast exact multiplication by the Hessian
Pearlmutter, B. A · 1994
Earlier work this paper cites.
ARPACK users’ guide: solution of large-scale eigenvalue problems with implicitly restarted Arnoldi methods
Lehoucq, R. B., Sorensen, D. C., and Yang, C · 1998
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I · 2000
Earlier work this paper cites.
On “natural” learning and pruning in multilayered perceptrons
Heskes, T · 2000
Earlier work this paper cites.
The ubiquitous Kronecker product
Loan, C. F · 2000
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Schraudolph, N. N · 2002
Earlier work this paper cites.
Gaussian processes for machine learning
Williams, C. K. and Rasmussen, C. E · 2006
Earlier work this paper cites.
An estimator for the diagonal of a matrix
Bekas, C., Kokiopoulou, E., and Saad, Y · 2007
Earlier work this paper cites.
LSMR: An iterative algorithm for sparse least-squares problems: Systems optimization laboratory
Fong, D. and Saunders, M · 2010
Earlier work this paper cites.
Deep learning via Hessian-free optimization
Martens, J · 2010
Earlier work this paper cites.
Finding Structure with Randomness: Probabilistic Algorithms for Constructing Approximate Matrix Decompositions
Halko, N., Martinsson, P. G., and Tropp, J. A · 2011
Earlier work this paper cites.
Training Deep and Recurrent Networks with Hessian-Free Optimization , pp. 479–535
Martens, J. and Sutskever, I · 2012
Earlier work this paper cites.
Matrix Computations
Golub, G. H. and Van Loan, C. F · 2013
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
An introduction to matrix concentration inequalities
Tropp, J. A · 2015
Earlier work this paper cites.
A kronecker-factored approximate Fisher matrix for convolution layers
Grosse, R. and Martens, J · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Tensor networks in a nutshell, 2017
Biamonte, J. and Bergholm, V · 2017
Earlier work this paper cites.
Practical Gauss-Newton optimisation for deep learning
Botev, A., Ritter, H., and Barber, D · 2017
Earlier work this paper cites.
Hand-waving and interpretive dance: an introductory course on tensor networks
Bridgeman, J. C. and Chubb, C. T · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Kipf, T. N. and Welling, M · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Cited alongside, same era.
Neumann optimizer: A practical optimization algorithm for deep neural networks, 2017
Krishnan, S., Xiao, Y., and Saurous, R. A · 2017
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond, 2017
Sagun, L., Bottou, L., and LeCun, Y · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Cited alongside, same era.
Exact natural gradient in deep linear networks and its application to the nonlinear case
Bernacchia, A., Lengyel, M., and Hennequin, G · 2018
Laplace redux - effortless bayesian deep learning
Daxberger, E., Kristiadi, A., Immer, A., Eschenhagen, R., Bauer, M., and Hennig, P · 2021
Later among the works it cites.
M-fac: Efficient matrix-free approximations of second-order information
Frantar, E., Kurtic, E., and Alistarh, D · 2021
Later among the works it cites.
Cockpit: A practical debugging tool for the training of deep neural networks
Schneider, F., Dangel, F., and Hennig, P · 2021
Later among the works it cites.
Analytic insights into structure and rank of neural network hessian maps
Singh, S. P., Bachmann, G., and Hofmann, T · 2021
Later among the works it cites.
Gradient descent on neurons and its link to approximate second-order optimization
Benzing, F · 2022
Later among the works it cites.
Efficient second-order optimization for neural networks with kernel machines
Chen, Y., Chen, Y., Chen, J., Wen, Z., and Huang, J · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., and Wanderman-Milne, S · 2018
Cited alongside, same era.
Gpytorch: Blackbox matrix-matrix gaussian process inference with gpu acceleration
Gardner, J., Pleiss, G., Weinberger, K. Q., Bindel, D., and Wilson, A. G · 2018
Cited alongside, same era.
Fast approximate natural gradient descent in a kronecker-factored eigenbasis, 2018
George, T., Laurent, C., Bouthillier, X., Ballas, N., and Vincent, P · 2018
Cited alongside, same era.
pytorch-hessian-eigenthings: efficient pytorch hessian eigendecomposition, 2018
Golmant, N., Yao, Z., Gholami, A., Mahoney, M., and Gonzalez, J · 2018
Cited alongside, same era.
Gradient descent happens in a tiny subspace, 2018
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Later among the works it cites.
ViViT: Curvature access through the generalized gauss-newton’s low-rank structure
Dangel, F., Tatzel, L., and Hennig, P · 2022
Later among the works it cites.
Monarch: Expressive structured matrices for efficient and accurate training
Dao, T., Chen, B., Sohoni, N. S., Desai, A., Poli, M., Grogan, J., Liu, A., Rao, A., Rudra, A., and Ré, C · 2022
Later among the works it cites.
The simplest, fastest repository for training/finetuning medium-sized gpts., 2022
Karpathy, A · 2022
Later among the works it cites.
Indiscriminate data poisoning attacks on neural networks
Lu, Y., Kamath, G., and Yu, Y · 2022
Later among the works it cites.
Merging models with fisher-weighted averaging
Matena, M. S. and Raffel, C. A · 2022
Later among the works it cites.
Kronecker-factored approximate curvature for modern neural network architectures
Eschenhagen, R., Immer, A., Turner, R. E., Schneider, F., and Hennig, P · 2023
Later among the works it cites.
Studying large language model generalization with influence functions, 2023
Grosse, R., Bae, J., Anil, C., Elhage, N., Tamkin, A., Tajdini, A., Steiner, B., Li, D., Durmus, E., Perez, E., Hubinger, E., Lukošiūtė, K., Nguyen, K., Joseph, N., McCandlish, S., Kaplan, J., and Bowman, S. R · 2023
Later among the works it cites.
Asdl: A unified interface for gradient preconditioning in pytorch, 2023
Osawa, K., Ishikawa, S., Yokota, R., Li, S., and Hoefler, T · 2023
Later among the works it cites.
ISAAC newton: Input-based approximate curvature for newton’s method
Petersen, F., Sutter, T., Borgelt, C., Huh, D., Kuehne, H., Sun, Y., and Deussen, O · 2023
Later among the works it cites.
CoLA: Exploiting Compositional Structure for Automatic and Efficient Numerical Linear Algebra
Potapczynski, A., Finzi, M., Pleiss, G., and Wilson, A. G · 2023
Later among the works it cites.
Lineax: unified linear solves and linear least-squares in jax and equinox
Rader, J., Lyons, T., and Kidger, P · 2023
Later among the works it cites.
How to guess a gradient
Singhal, U., Cheung, B., Chandra, K., Ragan-Kelley, J., Tenenbaum, J. B., Poggio, T. A., and Yu, S. X · 2023
Later among the works it cites.
Universality and sharp matrix concentration inequalities
Brailovskaya, T. and van Handel, R · 2024
Later among the works it cites.
How to compute hessian-vector products?
Dagréou, M., Ablin, P., Vaiter, S., and Moreau, T · 2024
Later among the works it cites.
Recent and upcoming developments in randomized numerical linear algebra for machine learning
Dereziński, M. and Mahoney, M. W · 2024
Later among the works it cites.
Xtrace: Making the most of every sample in stochastic trace estimation
Epperly, E. N., Tropp, J. A., and Webber, R. J · 2024
Later among the works it cites.
Skerch: Sketched matrix decompositions for PyTorch , 2024
Fernandez, A · 2024
Later among the works it cites.
Ginger: An efficient curvature approximation with linear complexity for general neural networks
Hao, Y., Cao, Y., and Mou, L · 2024
Later among the works it cites.
A sober look at LLMs for material discovery: Are they actually good for Bayesian optimization over molecules?
Kristiadi, A., Strieth-Kalthoff, F., Skreta, M., Poupart, P., Aspuru-Guzik, A., and Pleiss, G · 2024
Later among the works it cites.
Searching for efficient linear layers over a continuous space of structured matrices
Potapczynski, A., Qiu, S., Finzi, M. A., Ferri, C., Chen, Z., Goldblum, M., Bruss, C. B., Sa, C. D., and Wilson, A. G · 2024
Later among the works it cites.
Tradeoffs of diagonal fisher information matrix estimators
Soen, A. and Sun, K · 2024
Later among the works it cites.
Debiasing mini-batch quadratics for applications in deep learning, 2024
Tatzel, L., Mucsányi, B., Hackel, O., and Hennig, P · 2024
Later among the works it cites.
The LLM surgeon
van der Ouderaa, T. F. A., Nagel, M., van Baalen, M., Asano, Y. M., and Blankevoort, T · 2024
Later among the works it cites.
An improved empirical fisher approximation for natural gradient descent
Wu, X., Yu, W., Zhang, C., and Woodland, P · 2024
Later among the works it cites.
Bayesian low-rank adaptation for large language models
Yang, A. X., Robeyns, M., Wang, X., and Aitchison, L · 2024
Later among the works it cites.
Influence functions for scalable data attribution in diffusion models
Mlodozeniec, B., Eschenhagen, R., Bae, J., Immer, A., Krueger, D., and Turner, R · 2025
Closest in time.