Fetching the paper…
Reading the bibliography…
The existence of "lottery tickets" arXiv:1803.03635 at or near initialization raises the tantalizing question of whether large models are necessary in deep learning, or whether sparse networks can be quickly identified and trained without ever training the dense models that contain them.
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
Cover, T. M · 1965
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B. and Stork, D · 1992
Earlier work this paper cites.
Information theory and statistics: A tutorial
Csiszár, I. and Shields, P. C · 2004
Earlier work this paper cites.
Estimating mutual information
Kraskov, A., Stögbauer, H., and Grassberger, P · 2004
Earlier work this paper cites.
Pruning neural networks at initialization: Why are we missing the mark?
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M · 2009
Earlier work this paper cites.
Probability in Banach Spaces: Isoperimetry and Processes
Ledoux, M. and Talagrand, M · 2013
Earlier work this paper cites.
Lin, M., Chen, Q., and Yan, S · 2013
Earlier work this paper cites.
On the computational efficiency of training neural networks
Livni, R., Shalev-Shwartz, S., and Shamir, O · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Eie: Efficient inference engine on compressed deep neural network
Han, S., Liu, X., Mao, H., Pu, J., Pedram, A., Horowitz, M. A., and Dally, W. J · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2016
Earlier work this paper cites.
Probability in high dimensions
van Handel, R · 2016
Earlier work this paper cites.
Learning structured sparsity in deep neural networks
Wen, W., Wu, C., Wang, Y., Chen, Y., and Li, H · 2016
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Exploring the regularity of sparse structure in convolutional neural networks
Mao, H., Han, S., Pool, J., Li, W., Liu, X., Wang, Y., and Dally, W. J · 2017
Earlier work this paper cites.
Implicit regularization in deep learning
Neyshabur, B · 2017
Earlier work this paper cites.
Information-theoretic analysis of generalization capability of learning algorithms
Xu, A. and Raginsky, M · 2017
Earlier work this paper cites.
Chaining mutual information and tightening generalization bounds
Asadi, A., Abbe, E., and Verdú, S · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Cited alongside, same era.
Demystifying fixed
Gao, W., Oh, S., and Viswanath, P · 2018
Cited alongside, same era.
Amc: Automl for model compression and acceleration on mobile devices
He, Y., Lin, J., Liu, Z., Wang, H., Li, L.-J., and Han, S · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Snip: Single-shot network pruning based on connection sensitivity
Lee, N., Ajanthan, T., and Torr, P. H · 2018
Does learning require memorization? a short tale about a long tail
Feldman, V · 2020
Later among the works it cites.
Implicit regularization of random feature models
Jacot, A., Simsek, B., Spadaro, F., Hongler, C., and Gabriel, F · 2020
Later among the works it cites.
What’s hidden in a randomly weighted neural network?
Ramanujan, V., Wortsman, M., Kembhavi, A., Farhadi, A., and Rastegari, M · 2020
Later among the works it cites.
Pruning neural networks without any data by iteratively conserving synaptic flow
Tanaka, H., Kunin, D., Yamins, D. L., and Ganguli, S · 2020
Later among the works it cites.
Picking winning tickets before training by preserving gradient flow
Wang, C., Zhang, G., and Grosse, R · 2020
Later among the works it cites.
Sparch: Efficient architecture for sparse matrix multiplication
Zhang, Z., Wang, H., Han, S., and Dally, W. J · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rethinking the value of network pruning
Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T · 2018
Cited alongside, same era.
High-dimensional probability : an introduction with applications in data science
Vershynin, R · 2018
Cited alongside, same era.
To prune, or not to prune: Exploring the efficacy of pruning for model compression, 2018
Zhu, M. H. and Gupta, S · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Cited alongside, same era.
The difficulty of training sparse neural networks
Evci, U., Pedregosa, F., Gomez, A., and Elsen, E · 2019
Cited alongside, same era.
Stabilizing the lottery ticket hypothesis
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M · 2019
Cited alongside, same era.
Later among the works it cites.
A law of robustness for two-layers neural networks
Bubeck, S., Li, Y., and Nagaraj, D. M · 2021
Later among the works it cites.
Progressive skeletonization: Trimming more fat from a network at initialization
de Jorge, P., Sanyal, A., Behl, H., Torr, P., Rogez, G., and Dokania, P. K · 2021
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2021
Later among the works it cites.
Prospect pruning: Finding trainable weights at initialization using meta-gradients
Alizadeh, M., Tailor, S. A., Zintgraf, L. M., van Amersfoort, J., Farquhar, S., Lane, N. D., and Gal, Y · 2022
Later among the works it cites.
Unmasking the lottery ticket hypothesis: What’s encoded in a winning ticket’s mask?
Paul, M., Chen, F., Larsen, B. W., Frankle, J., Ganguli, S., and Dziugaite, G. K · 2022
Later among the works it cites.
Understanding pruning at initialization: An effective node-path balancing perspective
Pham, H., Ta, A., Liu, S., Le, D. D., and Tran-Thanh, L · 2022
Later among the works it cites.
Rare gems: Finding lottery tickets at initialization
Sreenivasan, K., Sohn, J.-y., Yang, L., Grinde, M., Nagle, A., Wang, H., Xing, E., Lee, K., and Papailiopoulos, D · 2022
Later among the works it cites.
Beyond the universal law of robustness: Sharper laws for random features and neural tangent kernels
Bombari, S., Kiyani, S., and Mondelli, M · 2023
Later among the works it cites.
A universal law of robustness via isoperimetry
Bubeck, S. and Sellke, M · 2023
Later among the works it cites.
Six lectures on linearized neural networks
Misiakiewicz, T. and Montanari, A · 2023
Later among the works it cites.
Simon, J. B., Karkada, D., Ghosh, N., and Belkin, M · 2023
Later among the works it cites.
Tractability from overparametrization: The example of the negative perceptron
Montanari, A., Zhong, Y., and Zhou, K · 2024
Closest in time.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Canatar, A., Bordelon, B., and Pehlevan, C · 2041
Closest in time.