Fetching the paper…
Reading the bibliography…
In the light of the fact that the stochastic gradient descent (SGD) often finds a flat minimum valley in the training loss, we propose a novel directional pruning method which searches for a sparse minimizer in or close to that flat region.
Semigroups of Linear Operators and Applications to Partial Differential Equations
A. Pazy · 1983
Earlier work this paper cites.
The asymptotic solution of linear differential systems, applications of the Levinson theorem
M. S. P. Eastham · 1989
Earlier work this paper cites.
Adaptive Algorithms and Stochastic Approximations
Albert Benveniste, Michel Métivier, and Pierre Priouret · 1990
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S. Denker, and Sara A. Solla · 1990
Earlier work this paper cites.
Weak convergence and local stability properties of fixed step size recursive algorithms
J. A. Bucklew, T. G. Kurtz, and W. A. Sethares · 1993
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
B. Hassibi, D. G. Stork, and G. J. Wolff · 1993
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David G. Stork · 1993
Earlier work this paper cites.
Asymptotic methods in the theory of Gaussian processes and fields
V. I. Piterbarg · 1996
Earlier work this paper cites.
Brownian Motion and Stochastic Calculus
Ioannis Karatzas and Steven Shreve · 1998
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and F.-F. Li · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Dual averaging methods for regularized stochastic learning and online optimization
Lin Xiao · 2010
Earlier work this paper cites.
Ordinary Differential Equations and Dynamical Systems
Gerald Teschl · 2012
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Song Han, Jeff Pool, John Tran, and William J. Dally · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
A generalized online mirror descent with applications to classification and regression
Francesco Orabona, Koby Crammer, and Nicolò Cesa-Bianchi · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J. Dally · 2016
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Variational dropout sparsifies deep neural networks
Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov · 2017
Cited alongside, same era.
Pruning convolutional neural networks for resource efficient inference
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz · 2017
Cited alongside, same era.
To prune, or not to prune: Exploring the efficacy of pruning for model compression
Michael H. Zhu and Suyog Gupta · 2017
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
Model compression and acceleration for deep neural networks: The principles, progress, and challenges
Y. Cheng, D. Wang, P. Zhou, and T. Zhang · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Later among the works it cites.
The difficulty of training sparse neural networks
Utku Evci, Fabian Pedregosa, Aidan Gomez, and Erich Elsen · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Later among the works it cites.
Stabilizing the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin · 2019
Later among the works it cites.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Cited alongside, same era.
When is a convolutional filter easy to learn?
Simon S. Du, Jason D. Lee, and Yuandong Tian · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Cited alongside, same era.
pytorch-hessian-eigentings: efficient pytorch hessian eigendecomposition, October 2018
Noah Golmant, Zhewei Yao, Amir Gholami, Michael Mahoney, and Joseph Gonzalez · 2018
Cited alongside, same era.
Using mode connectivity for loss landscape analysis
Akhilesh Gotmare, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Cited alongside, same era.
Gradient descent happens in a tiny subspace
Guy Gur-Ari, Daniel A. Roberts, and Ethan Dyer · 2018
Cited alongside, same era.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Later among the works it cites.
Asymmetric valleys: Beyond sharp and flat local minima
Haowei He, Gao Huang, and Yang Yuan · 2019
Later among the works it cites.
The implicit bias of gradient descent on nonseparable data
Ziwei Ji and Matus Telgarsky · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
Rethinking the value of network pruning
Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell · 2019
Later among the works it cites.
On connected sublevel sets in deep learning
Quynh Nguyen · 2019
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians
Vardan Papyan · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
How noise affects the hessian spectrum in overparameterized neural networks
Mingwei Wei and David J Schwab · 2019
Later among the works it cites.
Learning one-hidden-layer relu networks via gradient descent
Xiao Zhang, Yaodong Yu, Lingxiao Wang, and Quanquan Gu · 2019
Later among the works it cites.
10 breakthrough technologies, February 2020
2020
Closest in time.
Rigging the lottery: Making all tickets winners
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen · 2020
Closest in time.
AutoFIS: Automatic feature interaction selection in factorization models for click-through rate prediction
Bin Liu, Chenxu Zhu, Guilin Li, Weinan Zhang, Jincai Lai, Ruiming Tang, Xiuqiang He, Zhengguo Li, and Yong Yu · 2020
Closest in time.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Closest in time.