Fetching the paper…
Reading the bibliography…
Efficiently approximating local curvature information of the loss function is a key tool for optimization and compression of deep neural networks.
Inverting modified matrices
Max A Woodbury · 1950
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Optimal Brain Damage
Yann Le Cun, John S. Denker, and Sara A. Solla · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David G. Stork · 1992
Earlier work this paper cites.
A simple and effective method for removal of hidden units and weights
Masafumi Hagiwara · 1994
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-ichi Amari · 1998
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John C. Duchi, Elad Hazan, and Yoram Singer · 2010
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
New insights and perspectives on the natural gradient method
James Martens · 2014
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Diederik P Kingma, Tim Salimans, and Max Welling · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature, 2015
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Earlier work this paper cites.
Information geometry and its applications
Shun-ichi Amari · 2016
Earlier work this paper cites.
Distributed second-order optimization using kronecker-factored approximations
Jimmy Ba, Roger Grosse, and James Martens · 2016
Earlier work this paper cites.
A kronecker-factored approximate fisher matrix for convolution layers, 2016
Roger Grosse and James Martens · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Learning to prune deep neural networks via layer-wise optimal brain surgeon, 2017
Xin Dong, Shangyu Chen, and Sinno Jialin Pan · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications, 2017
Efficient full-matrix adaptive regularization
Naman Agarwal, Brian Bullins, Xinyi Chen, Elad Hazan, Karan Singh, Cyril Zhang, and Yi Zhang · 2019
Later among the works it cites.
The state of sparsity in deep neural networks, 2019
Trevor Gale, Erich Elsen, and Sara Hooker · 2019
Later among the works it cites.
Limitations of the empirical fisher approximation for natural gradient descent
Frederik Kunstner, Philipp Hennig, and Lukas Balles · 2019
Later among the works it cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Later among the works it cites.
Large-scale distributed second-order optimization using kronecker-factored approximate curvature for deep convolutional neural networks
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Cited alongside, same era.
Neumann optimizer: A practical optimization algorithm for deep neural networks
Shankar Krishnan, Ying Xiao, and Rif A Saurous · 2017
Cited alongside, same era.
A tutorial on fisher information, 2017
Alexander Ly, Maarten Marsman, Josine Verhagen, Raoul Grasman, and Eric-Jan Wagenmakers · 2017
Cited alongside, same era.
meProp: Sparsified back propagation for accelerated deep learning with reduced overfitting
Xu Sun, Xuancheng Ren, Shuming Ma, and Houfeng Wang · 2017
Cited alongside, same era.
Attention is all you need, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression, 2017
Michael Zhu and Suyog Gupta · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2018
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Well-read students learn better: On the importance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Eigendamage: Structured pruning in the kronecker-factored eigenbasis, 2019
Chaoqi Wang, Roger Grosse, Sanja Fidler, and Guodong Zhang · 2019
Later among the works it cites.
MLPrune: Multi-layer pruning for automated neural network compression, 2019
Wenyuan Zeng and Raquel Urtasun · 2019
Later among the works it cites.
Soft threshold weight reparameterization for learnable sparsity, 2020
Aditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman, Prateek Jain, Sham Kakade, and Ali Farhadi · 2020
Later among the works it cites.
Woodfisher: Efficient second-order approximation for neural network compression, 2020
Sidak Pal Singh and Dan Alistarh · 2020
Later among the works it cites.
On the interplay between noise and curvature and its effect on optimization and generalization
Valentin Thomas, Fabian Pedregosa, Bart Merriënboer, Pierre-Antoine Manzagol, Yoshua Bengio, and Nicolas Le Roux · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Later among the works it cites.
Adahessian: An adaptive second order optimizer for machine learning
Zhewei Yao, Amir Gholami, Sheng Shen, Kurt Keutzer, and Michael W Mahoney · 2020
Later among the works it cites.
M-FAC implementation: {https://github.com/IST-DASLab/M-FAC} , 2021
Elias Frantar and Eldar Kurtic · 2021
Closest in time.
Sparseml framework: https://github.com/neuralmagic/sparseml , 2021
Neural Magic Inc · 2021
Closest in time.