Fetching the paper…
Reading the bibliography…
The strong lottery ticket hypothesis holds the promise that pruning randomly initialized deep neural networks could offer a computationally efficient alternative to deep learning with stochastic gradient descent.
Optimal brain damage
Y. LeCun, J. S. Denker, and S. A. Solla · 1990
Earlier work this paper cites.
Generalization by weight-elimination with application to forecasting
A. Weigend, D. Rumelhart, and B. Huberman · 1991
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
B. Hassibi and D. G. Stork · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Exponentially small bounds on the expected optimum of the partition and subset sum problems
G. S. Lueker · 1998
Earlier work this paper cites.
Universal approximation using feedforward neural networks: A survey of some existing methods, and some new results
F. Scarselli and A. C. Tsoi · 1998
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
S. Srinivas and R. V. Babu · 2016
Earlier work this paper cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
D. Balduzzi, M. Frean, L. Leary, J. P. Lewis, K. W.-D. Ma, and B. McWilliams · 2017
Earlier work this paper cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
X. Dong, S. Chen, and S. J. Pan · 2017
Earlier work this paper cites.
Pruning filters for efficient convnets
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf · 2017
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
P. Molchanov, S. Tyree, T. Karras, T. Aila, and J. Kautz · 2017
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
J. Pennington, S. S. Schoenholz, and S. Ganguli · 2017
Cited alongside, same era.
Deep information propagation
S. S. Schoenholz, J. Gilmer, S. Ganguli, and J. Sohl-Dickstein · 2017
Cited alongside, same era.
The emergence of spectral universality in deep networks
J. Pennington, S. S. Schoenholz, and S. Ganguli · 2018
Cited alongside, same era.
Initialization of ReLUs for dynamical isometry
R. Burkholz and A. Dubatovka · 2019
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2019
Cited alongside, same era.
Logarithmic pruning is all you need
L. Orseau, M. Hutter, and O. Rivasplata · 2020
Later among the works it cites.
Optimal lottery tickets via subset sum: Logarithmic over-parameterization is sufficient
A. Pensia, S. Rajput, A. Nagle, H. Vishwakarma, and D. Papailiopoulos · 2020
Later among the works it cites.
Sanity-checking pruning methods: Random tickets can win the jackpot
J. Su, Y. Chen, T. Cai, T. Wu, R. Gao, L. Wang, and J. D. Lee · 2020
Later among the works it cites.
Pruning neural networks without any data by iteratively conserving synaptic flow
H. Tanaka, D. Kunin, D. L. Yamins, and S. Ganguli · 2020
Later among the works it cites.
Pruning via iterative ranking of sensitivity statistics, 2020
S. Verdenius, M. Stol, and P. Forré · 2020
Later among the works it cites.
Picking winning tickets before training by preserving gradient flow
C. Wang, G. Zhang, and R. B. Grosse · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep relu networks have surprisingly few activation patterns
B. Hanin and D. Rolnick · 2019
Cited alongside, same era.
Snip: single-shot network pruning based on connection sensitivity
N. Lee, T. Ajanthan, and P. H. S. Torr · 2019
Cited alongside, same era.
A mean field theory of batch normalization
G. Yang, J. Pennington, V. Rao, J. Sohl-Dickstein, and S. S. Schoenholz · 2019
Cited alongside, same era.
Deconstructing lottery tickets: Zeros, signs, and the supermask
H. Zhou, J. Lan, R. Liu, and J. Yosinski · 2019
Cited alongside, same era.
Linear mode connectivity and the lottery ticket hypothesis
J. Frankle, G. K. Dziugaite, D. Roy, and M. Carbin · 2020
Cited alongside, same era.
A signal propagation perspective for pruning neural networks at initialization
N. Lee, T. Ajanthan, S. Gould, and P. H. S. Torr · 2020
Cited alongside, same era.
Later among the works it cites.
Drawing early-bird tickets: Toward more efficient training of deep networks
H. You, C. Li, P. Xu, Y. Fu, Y. Wang, X. Chen, R. G. Baraniuk, Z. Wang, and Y. Lin · 2020
Later among the works it cites.
The elastic lottery ticket hypothesis
X. Chen, Y. Cheng, S. Wang, Z. Gan, J. Liu, and Z. Wang · 2021
Closest in time.
Robust pruning at initialization
S. Hayou, J.-F. Ton, A. Doucet, and Y. W. Teh · 2021
Closest in time.
Lottery ticket preserves weight correlation: Is it desirable or not?
N. Liu, G. Yuan, Z. Che, X. Shen, X. Ma, Q. Jin, J. Ren, J. Tang, S. Liu, and Y. Wang · 2021
Closest in time.
Sanity checks for lottery tickets: Does your winning ticket really win the jackpot?
X. Ma, G. Yuan, X. Shen, T. Chen, X. Chen, X. Chen, N. Liu, M. Qin, S. Liu, Z. Wang, and Y. Wang · 2021
Closest in time.
Validating the lottery ticket hypothesis with inertial manifold theory
Z. Zhang, J. Jin, Z. Zhang, Y. Zhou, X. Zhao, J. Ren, J. Liu, L. Wu, R. Jin, and D. Dou · 2021
Closest in time.
Plant ’n’ seek: Can you find the winning ticket?
J. Fischer and R. Burkholz · 2022
Closest in time.