Fetching the paper…
Reading the bibliography…
Model pruning is an essential procedure for building compact and computationally-efficient machine learning models.
One-shot pruning of recurrent neural networks by jacobian spectrum evaluation
Shunshi Zhang, M., and Stadie, B · 1912
Earlier work this paper cites.
Some inequalities for gaussian processes and applications
Gordon, Y · 1985
Earlier work this paper cites.
On Milman’s inequality and random subspaces which escape through a mesh in ℝ n \mathds{R}^{n}
Gordon, Y · 1988
Earlier work this paper cites.
Skeletonization: A technique for trimming the fat from a network via relevance assessment
Mozer, M. C., and Smolensky, P · 1989
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B., and Stork, D. G · 1993
Earlier work this paper cites.
Optimal brain surgeon: Extensions and performance comparisons
Hassibi, B., Stork, D. G., and Wolff, G · 1994
Earlier work this paper cites.
On the kalman filtering method in neural network training and pruning
Sum, J., Leung, C.-S., Young, G. H., and Kan, W.-K · 1999
Earlier work this paper cites.
Analysis of local appearance-based face recognition: Effects of feature selection and feature normalization
Ekenel, H. K., and Stiefelhagen, R · 2006
Earlier work this paper cites.
Message-passing algorithms for compressed sensing
Donoho, D. L., Maleki, A., and Montanari, A · 2009
Earlier work this paper cites.
Statistical normalization and back propagation for classification
Jayalakshmi, T., and Santhakumaran, A · 2011
Earlier work this paper cites.
Accurate prediction of phase transitions in compressed sensing via a connection to minimax denoising
Donoho, D. L., Johnstone, I., and Montanari, A · 2013
Earlier work this paper cites.
The squared-error of generalized lasso: A precise analysis
Oymak, S., Thrampoulidis, C., and Hassibi, B · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N · 2014
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S., and Szegedy, C · 2015
Earlier work this paper cites.
Lasso with non-linear measurements is equivalent to one with linear measurements
Thrampoulidis, C., Abbasi, E., and Hassibi, B · 2015
Earlier work this paper cites.
Regularized linear regression: A precise analysis of the estimation error
Thrampoulidis, C., Oymak, S., and Hassibi, B · 2015
Earlier work this paper cites.
Training skinny deep neural networks with iterative hard thresholding methods
Jin, X., Yuan, X., Feng, J., and Yan, S · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Earlier work this paper cites.
Net-trim: Convex pruning of deep neural networks with performance guarantee
Aghasi, A., Abdi, A., Nguyen, N., and Romberg, J · 2017
Cited alongside, same era.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Dong, X., Chen, S., and Pan, S · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B. E., Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2017
Cited alongside, same era.
Efficient processing of deep neural networks: A tutorial and survey
Sze, V., Chen, Y.-H., Yang, T.-J., and Emer, J. S · 2017
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Arora, S., Cohen, N., and Hazan, E · 2018
Cited alongside, same era.
To understand deep learning we need to understand kernel learning
Belkin, M., Ma, S., and Mandal, S · 2018
Stochastic gradient/mirror descent: Minimax optimality and implicit regularization
Azizan, N., and Hassibi, B · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
Later among the works it cites.
Two models of double descent for weak features
Belkin, M., Hsu, D., and Xu, J · 2019
Later among the works it cites.
Does data interpolation contradict statistical optimality?
Belkin, M., Rakhlin, A., and Tsybakov, A. B · 2019
Later among the works it cites.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S. S., Lee, J. D., Li, H., Wang, L., and Zhai, X · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Risk and parameter convergence of logistic regression
Ji, Z., and Telgarsky, M · 2018
Cited alongside, same era.
Snip: Single-shot network pruning based on connection sensitivity
Lee, N., Ajanthan, T., and Torr, P. H · 2018
Cited alongside, same era.
Just interpolate: Kernel" ridgeless" regression can generalize
Liang, T., and Rakhlin, A · 2018
Cited alongside, same era.
Frankle, J., and Carbin, M · 2019
Later among the works it cites.
The lottery ticket hypothesis at scale
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M · 2019
Later among the works it cites.
Path length bounds for gradient descent and flow
Gupta, C., Balakrishnan, S., and Ramdas, A · 2019
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J · 2019
Later among the works it cites.
Li, M., Soltanolkotabi, M., and Oymak, S · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei, S., and Montanari, A · 2019
Later among the works it cites.
Montanari, A., Ruan, F., Sohn, Y., and Yan, J · 2019
Later among the works it cites.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
Nacson, M. S., Srebro, N., and Soudry, D · 2019
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2019
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Oymak, S., and Soltanolkotabi, M · 2019
Later among the works it cites.
Luck matters: Understanding training dynamics of deep relu networks
Tian, Y., Jiang, T., Gong, Q., and Morcos, A · 2019
Later among the works it cites.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Zhou, H., Lan, J., Liu, R., and Yosinski, J · 2019
Later among the works it cites.
A model of double descent for high-dimensional logistic regression
Deng, Z., Kammoun, A., and Thrampoulidis, C · 2020
Closest in time.
Proving the lottery ticket hypothesis: Pruning is all you need
Malach, E., Yehudai, G., Shalev-Shwartz, S., and Shamir, O · 2020
Closest in time.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
Oymak, S., and Soltanolkotabi, M · 2020
Closest in time.
Picking winning tickets before training by preserving gradient flow
Wang, C., Zhang, G., and Grosse, R · 2020
Closest in time.