Fetching the paper…
Reading the bibliography…
Pruning deep neural networks is a widely used strategy to alleviate the computational burden in machine learning.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla · 1989
Earlier work this paper cites.
Linear and nonlinear extension of the pseudo-inverse solution for learning boolean functions
F Vallet, J-G Cailton, and Ph Refregier · 1989
Earlier work this paper cites.
On the ability of the optimal perceptron to generalise
M Opper, W Kinzel, J Kleinz, and R Nehl · 1990
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
Stuart Geman, Elie Bienenstock, and René Doursat · 1992
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David Stork · 1992
Earlier work this paper cites.
Statistical mechanics of learning from examples
Hyunjune Sebastian Seung, Haim Sompolinsky, and Naftali Tishby · 1992
Earlier work this paper cites.
Why lottery ticket wins? a theoretical perspective of sample complexity on sparse neural networks
Shuai Zhang, Meng Wang, Sijia Liu, Pin-Yu Chen, and Jinjun Xiong · 1992
Earlier work this paper cites.
Reflections after refereeing papers for nips
Leo Breiman · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
The optimal brain surgeon for pruning neural network architecture applied to multivariate calibration
RJ Poppi and DL Massart · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Measuring the intrinsic dimension of objective landscapes
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Cited alongside, same era.
Stabilizing the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M Roy, and Michael Carbin · 2019
Cited alongside, same era.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker · 2019
Cited alongside, same era.
Towards understanding the role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2019
Cited alongside, same era.
A jamming transition from under-to over-parametrization affects generalization in deep learning
S Spigler, M Geiger, S d’Ascoli, L Sagun, G Biroli, and M Wyart · 2019
Cited alongside, same era.
What causes the test error? going beyond bias-variance via anova
Licong Lin and Edgar Dobriban · 2021
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Later among the works it cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
Curtis G Northcutt, Anish Athalye, and Jonas Mueller · 2021
Later among the works it cites.
Franco Pellegrini and Giulio Biroli · 2021
Later among the works it cites.
The intrinsic dimension of images and its impact on learning
Phil Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum, and Tom Goldstein · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
High-dimensional dynamics of generalization error in neural networks
Madhu S. Advani, Andrew M. Saxe, and Haim Sompolinsky · 2020
Cited alongside, same era.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2020
Cited alongside, same era.
What is the state of neural network pruning?
Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag · 2020
Cited alongside, same era.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin · 2020
Cited alongside, same era.
A brief prehistory of double descent
Marco Loog, Tom Viering, Alexander Mey, Jesse H Krijthe, and David MJ Tax · 2020
Cited alongside, same era.
Proving the lottery ticket hypothesis: Pruning is all you need
Eran Malach, Gilad Yehudai, Shai Shalev-Schwartz, and Ohad Shamir · 2020
Cited alongside, same era.
Logarithmic pruning is all you need
Laurent Orseau, Marcus Hutter, and Omar Rivasplata · 2020
Cited alongside, same era.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborová · 2021
Later among the works it cites.
The effect of data dimensionality on neural network prunability
Zachary Ankner, Alex Renda, Gintare Karolina Dziugaite, Jonathan Frankle, and Tian Jin · 2022
Later among the works it cites.
Most activation functions can win the lottery without excessive depth
Rebekka Burkholz · 2022
Later among the works it cites.
Proving the strong lottery ticket hypothesis for convolutional neural networks
Arthur da Cunha, Emanuele Natale, and Laurent Viennot · 2022
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2022
Later among the works it cites.
Sparse double descent: Where network pruning aggravates overfitting
Zheng He, Zeke Xie, Quanzhi Zhu, and Zengchang Qin · 2022
Later among the works it cites.
Pruning’s effect on generalization through the lens of training and regularization
Tian Jin, Michael Carbin, Daniel M. Roy, Jonathan Frankle, and Gintare Karolina Dziugaite · 2022
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2022
Later among the works it cites.
Unmasking the lottery ticket hypothesis: What’s encoded in a winning ticket’s mask?
Mansheej Paul, Feng Chen, Brett W Larsen, Jonathan Frankle, Surya Ganguli, and Gintare Karolina Dziugaite · 2022
Later among the works it cites.
Analyzing lottery ticket hypothesis from PAC-bayesian theory perspective
Keitaro Sakamoto and Issei Sato · 2022
Later among the works it cites.
Calibrating the rigged lottery: Making all tickets reliable
Bowen Lei, Ruqi Zhang, Dongkuan Xu, and Bani Mallick · 2023
Closest in time.
Theoretical characterization of how neural network pruning affects its generalization
Hongru Yang, Yingbin Liang, Xiaojie Guo, Lingfei Wu, and Zhangyang Wang · 2023
Closest in time.