Fetching the paper…
Reading the bibliography…
We present MLRG Deep Curvature suite, a PyTorch-based, open-source package for analysis and visualisation of neural network curvature and loss landscape.
Distribution of eigenvalues for some sets of random matrices
Marchenko, V. A. and Pastur, L. A · 1967
Earlier work this paper cites.
A stochastic estimator of the trace of the influence matrix for Laplacian smoothing splines
Hutchinson, M. F · 1990
Earlier work this paper cites.
Characteristic vectors of bordered matrices with infinite dimensions i
Wigner, E. P · 1993
Earlier work this paper cites.
Matrices, moments and quadrature
Golub, G. H. and Meurant, G · 1994
Earlier work this paper cites.
Fast exact multiplication by the Hessian
Pearlmutter, B. A · 1994
Earlier work this paper cites.
Some large-scale matrix computation problems
Bai, Z., Fahey, G., and Golub, G · 1996
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C. M · 2006
Earlier work this paper cites.
The Lanczos and conjugate gradient algorithms in finite precision arithmetic
Meurant, G. and Strakoš, Z · 2006
Earlier work this paper cites.
The Oxford handbook of random matrix theory
Akemann, G., Baik, J., and Di Francesco, P · 2011
Earlier work this paper cites.
Matrix computations , volume 3
Golub, G. H. and Van Loan, C. F · 2012
Earlier work this paper cites.
Training deep and recurrent networks with Hessian-free optimization
Martens, J. and Sutskever, I · 2012
Earlier work this paper cites.
Topics in random matrix theory , volume 132
Tao, T · 2012
Earlier work this paper cites.
Krylov subspace descent for deep learning
Vinyals, O. and Povey, D · 2012
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Nesterov, Y · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Cited alongside, same era.
Entropy-SGD: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2016
Gpytorch: Blackbox matrix-matrix Gaussian process inference with GPU acceleration
Gardner, J., Pleiss, G., Weinberger, K. Q., Bindel, D., and Wilson, A. G · 2018
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G · 2018
Later among the works it cites.
Kohler, J., Daneshmand, H., Lucchi, A., Zhou, M., Neymeyr, K., and Hofmann, T · 2018
Later among the works it cites.
A scalable laplace approximation for neural networks
Ritter, H., Botev, A., and Barber, D · 2018
Later among the works it cites.
How does batch normalization help optimization?
Santurkar, S., Tsipras, D., Ilyas, A., and Madry, A · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Approximating spectral densities of large matrices
Lin, L., Saad, Y., and Yang, C · 2016
Cited alongside, same era.
Second-order optimization for neural networks
Martens, J · 2016
Cited alongside, same era.
Eigenvalues of the Hessian in deep learning: Singularity and beyond
Sagun, L., Bottou, L., and LeCun, Y · 2016
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., and Goldstein, T · 2017
Cited alongside, same era.
Automatic differentiation in Pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
Pennington, J. and Bahri, Y · 2017
Cited alongside, same era.
Yao, Z., Gholami, A., Lei, Q., Keutzer, K., and Mahoney, M. W · 2018
Later among the works it cites.
Backpack: Packing more into backprop
Dangel, F., Kunstner, F., and Hennig, P · 2019
Closest in time.
An investigation into neural net optimization via Hessian eigenvalue density
Ghorbani, B., Krishnan, S., and Xiao, Y · 2019
Closest in time.
Meme: An accurate maximum entropy method for efficient approximations in large-scale machine learning
Granziol, D., Ru, B., Zohren, S., Dong, X., Osborne, M., and Roberts, S · 2019
Closest in time.
Asymmetric valleys: Beyond sharp and flat local minima
He, H., Huang, G., and Yuan, Y · 2019
Closest in time.
Subspace inference for bayesian deep learning
Izmailov, P., Maddox, W. J., Kirichenko, P., Garipov, T., Vetrov, D., and Wilson, A. G · 2019
Closest in time.
A simple baseline for bayesian uncertainty in deep learning
Maddox, W. J., Izmailov, P., Garipov, T., Vetrov, D. P., and Wilson, A. G · 2019
Closest in time.
Towards understanding the true loss surface of deep neural networks using random matrix theory and iterative spectral methods
Anonymous · 2020
Closest in time.