Fetching the paper…
Reading the bibliography…
Many existing Neural Network pruning approaches rely on either retraining or inducing a strong bias in order to converge to a sparse solution throughout training.
An algorithm for quadratic programming
Frank, M., Wolfe, P., et al · 1956
Earlier work this paper cites.
Constrained minimization methods
Levitin, E. S. and Polyak, B. T · 1966
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1989
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B. and Stork, D · 1992
Earlier work this paper cites.
Characterization of the subdifferential of some matrix norms
Watson, G. A · 1992
Earlier work this paper cites.
Model selection and estimation in regression with grouped variables
Yuan, M. and Lin, Y · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
A singular value thresholding algorithm for matrix completion
Cai, J.-F., Candès, E. J., and Shen, Z · 2010
Earlier work this paper cites.
Sparse prediction with the k k -support norm
Argyriou, A., Foygel, R., and Srebro, N · 2012
Earlier work this paper cites.
Revisiting frank-wolfe: Projection-free sparse convex optimization
Jaggi, M · 2013
Earlier work this paper cites.
Block-coordinate frank-wolfe optimization for structural svms
Lacoste-Julien, S., Jaggi, M., Schmidt, M., and Pletscher, P · 2013
Earlier work this paper cites.
Exploiting linear structure within convolutional networks for efficient evaluation
Denton, E. L., Zaremba, W., Bruna, J., LeCun, Y., and Fergus, R · 2014
Earlier work this paper cites.
Speeding-up convolutional neural networks using fine-tuned cp-decomposition
Lebedev, V., Ganin, Y., Rakhuba, M., Oseledets, I., and Lempitsky, V · 2014
Earlier work this paper cites.
The ordered weighted ℓ 1 \ell_{1} norm: Atomic formulation, projections, and algorithms
Zeng, X. and Figueiredo, M. A. T · 2014
Earlier work this paper cites.
Block power method for svd decomposition
Bentbib, A. and Kanber, A · 2015
Earlier work this paper cites.
Fast and scalable lasso via stochastic frank-wolfe methods with a convergence guarantee
Frandi, E., Nanculef, R., Lodi, S., Sartori, C., and Suykens, J. A. K · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Compression of deep convolutional neural networks for fast and low power mobile applications
Kim, Y.-D., Park, E., Yoo, S., Choi, T., Yang, L., and Shin, D · 2015
Earlier work this paper cites.
Tiny imagenet visual recognition challenge
Le, Y. and Yang, X · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Convolutional neural networks with low-rank regularization
Tai, C., Xiao, T., Zhang, Y., Wang, X., and E, W · 2015
Earlier work this paper cites.
Accelerating very deep convolutional networks for classification and detection
Zhang, X., Zou, J., He, K., and Sun, J · 2015
Earlier work this paper cites.
Lazysvd: Even faster svd decomposition yet without agonizing pain
Allen-Zhu, Z. and Li, Y · 2016
Earlier work this paper cites.
Learning the number of neurons in deep networks
Alvarez, J. M. and Salzmann, M · 2016
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B · 2016
Earlier work this paper cites.
Variance-reduced and projection-free stochastic optimization
Hazan, E. and Luo, H · 2016
Earlier work this paper cites.
Convergence rate of frank-wolfe for non-convex objectives
Lacoste-Julien, S · 2016
Earlier work this paper cites.
Pruning filters for efficient convnets
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P · 2016
Earlier work this paper cites.
Fitting spectral decay with the k-support norm
McDonald, A., Pontil, M., and Stamos, D · 2016
Cited alongside, same era.
Stochastic frank-wolfe methods for nonconvex optimization
Reddi, S. J., Sra, S., Póczos, B., and Smola, A · 2016
Cited alongside, same era.
Learning structured sparsity in deep neural networks
Wen, W., Wu, C., Wang, Y., Chen, Y., and Li, H · 2016
Cited alongside, same era.
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Pyramid scene parsing network
Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J · 2016
Cited alongside, same era.
Compression-aware training of deep networks
Alvarez, J. M. and Salzmann, M · 2017
Cited alongside, same era.
The generalization-stability tradeoff in neural network pruning
Bartoldson, B., Morcos, A., Barbu, A., and Erlebacher, G · 2020
Later among the works it cites.
Experiment tracking with weights and biases, 2020
Biewald, L · 2020
Later among the works it cites.
What is the state of neural network pruning?
Blalock, D., Gonzalez Ortiz, J. J., Frankle, J., and Guttag, J · 2020
Later among the works it cites.
Once-for-all: Train one network and specialize it for efficient deployment
Cai, H., Gan, C., Wang, T., Zhang, Z., and Han, S · 2020
Later among the works it cites.
Projection-free adaptive gradients for large-scale optimization
Combettes, C. W., Spiegel, C., and Pokutta, S · 2020
Later among the works it cites.
Rigging the lottery: Making all tickets winners
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The group k-support norm for learning with structured sparsity
Rao, N., Dudík, M., and Harchaoui, Z · 2017
Cited alongside, same era.
L2 regularization versus batch and weight normalization
van Laarhoven, T · 2017
Cited alongside, same era.
Coordinating filters for faster deep neural networks
Wen, W., Xu, C., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Cited alongside, same era.
To prune, or not to prune: Exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Cited alongside, same era.
Deep frank-wolfe for neural network optimization
Berrada, L., Zisserman, A., and Kumar, M. P · 2018
Cited alongside, same era.
Learning-compression algorithms for neural net pruning
Carreira-Perpinán, M. A. and Idelbayev, Y · 2018
Cited alongside, same era.
Low-rank compression of neural nets: Learning the rank of each layer
Idelbayev, Y. and Carreira-Perpinán, M. A · 2020
Later among the works it cites.
Position-based scaled gradient for model quantization and pruning
Kim, J., Yoo, K., and Kwak, N · 2020
Later among the works it cites.
Soft threshold weight reparameterization for learnable sparsity
Kusupati, A., Ramanujan, V., Somani, R., Wortsman, M., Jain, P., Kakade, S., and Farhadi, A · 2020
Later among the works it cites.
Layer-adaptive sparsity for the magnitude-based pruning
Lee, J., Park, S., Mo, S., Ahn, S., and Shin, J · 2020
Later among the works it cites.
Dynamic model pruning with feedback
Lin, T., Stich, S. U., Barba, L., Dmitriev, D., and Jaggi, M · 2020
Later among the works it cites.
Dynamic sparse training: Find efficient sparse network from scratch with trainable masked layers
Liu, J., Xu, Z., Shi, R., Cheung, R. C. C., and So, H. K · 2020
Later among the works it cites.
Stochastic frank-wolfe for constrained finite-sum minimization
Négiar, G., Dresdner, G., Tsai, A., El Ghaoui, L., Locatello, F., Freund, R., and Pedregosa, F · 2020
Later among the works it cites.
Deep neural network training with frank-wolfe
Pokutta, S., Spiegel, C., and Zimmer, M · 2020
Later among the works it cites.
Comparing rewinding and fine-tuning in neural network pruning
Renda, A., Frankle, J., and Carbin, M · 2020
Later among the works it cites.
On frank-wolfe optimization for adversarial robustness and interpretability
Tsiligkaridis, T. and Roberts, J · 2020
Later among the works it cites.
Trp: Trained rank pruning for efficient deep neural networks
Xu, Y., Li, Y., Zhang, S., Wen, W., Wang, B., Qi, Y., Chen, Y., Lin, W., and Xiong, H · 2020
Later among the works it cites.
Learning low-rank deep neural networks via singular vector orthogonality regularization and singular value sparsification
Yang, H., Tang, M., Wen, W., Yan, F., Hu, D., Li, A., Li, H., and Chen, Y · 2020
Later among the works it cites.
Neuron-level structured pruning using polarization regularizer
Zhuang, T., Zhang, Z., Huang, Y., Zeng, X., Shuang, K., and Li, X · 2020
Later among the works it cites.
How I Learned To Stop Worrying And Love Retraining
Zimmer, M., Spiegel, C., and Pokutta, S · 2020
Later among the works it cites.
Cindy: Conditional gradient-based identification of non-linear dynamics – noise-robust recovery
Carderera, A., Pokutta, S., Schütte, C., and Weiser, M · 2021
Later among the works it cites.
Complexity of linear minimization and projection on some sets
Combettes, C. W. and Pokutta, S · 2021
Later among the works it cites.
Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A · 2021
Later among the works it cites.
Network pruning that matters: A case study on retraining variants
Le, D. H. and Hua, B.-S · 2021
Later among the works it cites.
Compressing neural networks: Towards determining the optimal layer-wise decomposition
Liebenwein, L., Maalouf, A., Gal, O., Feldman, D., and Rus, D · 2021
Later among the works it cites.
Growing efficient deep networks by structured continuous sparsification
Yuan, X., Savarese, P. H. P., and Maire, M · 2021
Later among the works it cites.
Conditional gradient methods
Braun, G., Carderera, A., Combettes, C. W., Hassani, H., Karbasi, A., Mokhtari, A., and Pokutta, S · 2022
Closest in time.
Learning pruning-friendly networks via frank-wolfe: One-shot, any-sparsity, and no retraining
Miao, L., Luo, X., Chen, T., Chen, W., Liu, D., and Wang, Z · 2022
Closest in time.
Cram: A compression-aware minimizer
Peste, A., Vladu, A., Alistarh, D., and Lampert, C. H · 2022
Closest in time.