Fetching the paper…
Reading the bibliography…
The recent focus on the efficiency of deep neural networks (DNNs) has led to significant work on model compression approaches, of which weight pruning is one of the most popular.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1990
Earlier work this paper cites.
A simple and effective method for removal of hidden units and weights
Hagiwara, M · 1994
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Torchvision the machine-vision package of torch
Marcel, S. and Rodriguez, Y · 2010
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Han, S., Pool, J., Tran, J., and Dally, W. J · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Pruning filters for efficient convnets
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Channel pruning for accelerating very deep neural networks
He, Y., Zhang, X., and Sun, J · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H · 2017
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Earlier work this paper cites.
N2N learning: Network to network compression via policy gradient reinforcement learning
Ashok, A., Rhinehart, N., Beainy, F., and Kitani, K. M · 2018
Earlier work this paper cites.
ProxylessNAS: Direct neural architecture search on target task and hardware
Cai, H., Zhu, L., and Han, S · 2018
Earlier work this paper cites.
Mean replacement pruning
Evci, U., Le Roux, N., Castro, P., and Bottou, L · 2018
Earlier work this paper cites.
AMC: AutoML for model compression and acceleration on mobile devices
He, Y., Lin, J., Liu, Z., Wang, H., Li, L.-J., and Han, S · 2018
Earlier work this paper cites.
Once-for-All: Train one network and specialize it for efficient deployment
Cai, H., Gan, C., Wang, T., Zhang, Z., and Han, S · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2019
Cited alongside, same era.
Filter pruning via geometric median for deep convolutional neural networks acceleration
He, Y., Liu, P., Wang, Z., Hu, Z., and Yang, Y · 2019
Cited alongside, same era.
Knapsack pruning with inner distillation
Aflalo, Y., Noy, A., Lin, M., Friedman, I., and Zelnik, L · 2020
Cited alongside, same era.
Fast sparse convnets
Elsen, E., Dukhan, M., Gale, T., and Simonyan, K · 2020
Cited alongside, same era.
Rigging the lottery: Making all tickets winners
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Cited alongside, same era.
Sparse GPU kernels for deep learning
Gale, T., Zaharia, M., Young, C., and Elsen, E · 2020
Top-KAST: Top-K always sparse training
Jayakumar, S. M., Pascanu, R., Rae, J. W., Osindero, S., and Elsen, E · 2021
Later among the works it cites.
Block pruning for faster transformers
Lagunas, F., Charlaix, E., Sanh, V., and Rush, A · 2021
Later among the works it cites.
BRECQ: Pushing the limit of post-training quantization by block reconstruction
Li, Y., Gong, R., Tan, X., Yang, Y., Hu, P., Zhang, Q., Yu, F., Wang, W., and Gu, S · 2021
Later among the works it cites.
Compressing neural networks: Towards determining the optimal layer-wise decomposition
Liebenwein, L., Maalouf, A., Feldman, D., and Rus, D · 2021
Later among the works it cites.
Group fisher pruning for practical network compression
Liu, L., Zhang, S., Kuang, Z., Zhou, A., Xue, J.-H., Wang, X., Chen, Y., Yang, W., Liao, Q., and Zhang, W · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Single path one-shot neural architecture search with uniform sampling
Guo, Z., Zhang, X., Mu, H., Heng, W., Liu, Z., Wei, Y., and Sun, J · 2020
Cited alongside, same era.
Inducing and exploiting activation sparsity for fast inference on deep neural networks
Kurtz, M., Kopinsky, J., Gelashvili, R., Matveev, A., Carr, J., Goin, M., Leiserson, W., Moore, S., Nell, B., Shavit, N., and Alistarh, D · 2020
Cited alongside, same era.
Soft threshold weight reparameterization for learnable sparsity
Kusupati, A., Ramanujan, V., Somani, R., Wortsman, M., Jain, P., Kakade, S., and Farhadi, A · 2020
Cited alongside, same era.
Up or down? adaptive rounding for post-training quantization
Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T · 2020
Cited alongside, same era.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V., Wolf, T., and Rush, A. M · 2020
Cited alongside, same era.
Woodfisher: Efficient second-order approximation for neural network compression
Singh, S. P. and Alistarh, D · 2020
Cited alongside, same era.
Markov, I., Ramezani, H., and Alistarh, D · 2021
Later among the works it cites.
Accelerating sparse deep neural networks
Mishra, A., Latorre, J. A., Pool, J., Stosic, D., Stosic, D., Venkatesh, G., Yu, C., and Micikevicius, P · 2021
Later among the works it cites.
A white paper on neural network quantization
Nagel, M., Fournarakis, M., Amjad, R. A., Bondarenko, Y., van Baalen, M., and Blankevoort, T · 2021
Later among the works it cites.
DeepSparse, 2021
NeuralMagic · 2021
Later among the works it cites.
Ac/dc: Alternating compressed/decompressed training of deep neural networks
Peste, A., Iofinova, E., Vladu, A., and Alistarh, D · 2021
Later among the works it cites.
Channel permutations for N:M sparsity
Pool, J. and Yu, C · 2021
Later among the works it cites.
Powerpropagation: A sparsity inducing weight reparameterisation
Schwarz, J., Jayakumar, S., Pascanu, R., Latham, P., and Teh, Y · 2021
Later among the works it cites.
NetAdaptV2: Efficient neural architecture search with fast super-network training and architecture optimization
Yang, T.-J., Liao, Y.-L., and Sze, V · 2021
Later among the works it cites.
HAWQ-v3: Dyadic neural network quantization
Yao, Z., Dong, Z., Zheng, Z., Gholami, A., Yu, J., Tan, E., Wang, L., Huang, Q., Wang, Y., Mahoney, M., et al · 2021
Later among the works it cites.
Learning N:M fine-grained structured sparse neural networks from scratch
Zhou, A., Ma, Y., Zhu, J., Liu, J., Zhang, Z., Yuan, K., Sun, W., and Li, H · 2021
Later among the works it cites.
YOLOv5, 2022
Jocher, G · 2022
Closest in time.
The Optimal BERT Surgeon: Scalable and accurate second-order pruning for large language models
Kurtic, E., Campos, D., Nguyen, T., Frantar, E., Kurtz, M., Fineran, B., Goin, M., and Alistarh, D · 2022
Closest in time.