Fetching the paper…
Reading the bibliography…
Modern deep neural networks, particularly recent large language models, come with massive model sizes that require significant computational and storage resources.
L. Yao, R. Pi, H. Xu, W. Zhang, Z. Li, and T. Zhang, “Joint-DetNAS: Upgrade your detector with NAS, pruning and dynamic distillation,” in
1908
Earlier work this paper cites.
S. Hanson and L. Pratt, “Comparing biases for minimal network construction with back-propagation,” in
1988
Earlier work this paper cites.
Y. LeCun, J. Denker, and S. Solla, “Optimal brain damage,” in
1989
Earlier work this paper cites.
B. Hassibi and D. G. Stork, “Second order derivatives for network pruning: Optimal brain surgeon,” in
1992
Earlier work this paper cites.
R. Reed, “Pruning algorithms-a survey,”
1993
Earlier work this paper cites.
M. Marcus, B. Santorini, and M. A. Marcinkiewicz, “Building a large annotated corpus of English: The penn treebank,”
1993
Earlier work this paper cites.
R. Tibshirani, “Regression shrinkage and selection via the lasso,”
1996
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: a method for automatic evaluation of machine translation,” in
2002
Earlier work this paper cites.
F. J. Huang and Y. LeCun, “Learning methods for generic object recognition with invariance to pose and lighting,” in
2004
Earlier work this paper cites.
Z. Wang, A. Bovik, H. Sheikh, and E. Simoncelli, “Image quality assessment: from error visibility to structural similarity,”
2004
Earlier work this paper cites.
Z. Wang, A. Bovik, H. Sheikh, and E. Simoncelli, “Image quality assessment: from error visibility to structural similarity,”
2004
Earlier work this paper cites.
M. Yuan and Y. Lin, “Model selection and estimation in regression with grouped variables,”
2006
Earlier work this paper cites.
R. Hadsell, S. Chopra, and Y. LeCun, “Dimensionality reduction by learning an invariant mapping,” in
2006
Earlier work this paper cites.
P. Fletcher, S. Venkatasubramanian, and S. Joshi, “Robust statistics on riemannian manifolds via the geometric median,” in
2008
Earlier work this paper cites.
M. Everingham, C. Luc Van Gool, J. Winn, and A. Zisserman, “The PASCAL visual object classes (VOC) challenge 2007,”
2008
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,”
2009
Earlier work this paper cites.
S. J. Pan and Q. Yang, “A suvery on tranfer learning,”
2009
Earlier work this paper cites.
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein, “Distributed optimization and statistical learning via the alternating direction method of multipliers,”
2010
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Esteve, “TED-LIUM: an automatic speech recognition dedicated corpus,” in
2012
Earlier work this paper cites.
I. Sutskever, J. Martens, G. Dahl, and G. Hinton, “On the importance of initialization and momentum in deep learning,” in
2013
Earlier work this paper cites.
R. Socher, A. Perelygin, and J. Wu, et al., “Recursive deep models for semantic compositionality over a sentiment treebank,” in
2013
Earlier work this paper cites.
E. Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus, “Exploiting linear structure within convolutional networks for efficient evaluation,” in
2014
Earlier work this paper cites.
A. M. Saxe, J. L. McClelland, and S. Ganguli, “Exact solutions to the nonlinear dynamics of learning in deep linear neural networks,” in
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and L. Zitnick, “Microsoft COCO: Common objects in context,” in
2014
Earlier work this paper cites.
P. Young, A. Lai, and M. Hodosh, et al., “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” in
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in
2015
Earlier work this paper cites.
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural network,” in
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, and H. Su, et al., “ImageNet large scale visual recognition challenge,”
2015
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,”
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in
2015
Earlier work this paper cites.
Z. Liu, P. Luo, X. Wang, and X. Tang, “Deep learning face attributes in the wild,” in
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, , and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,” in
2016
Earlier work this paper cites.
W. Wen, C. Wu, Y. Wang, Y. Chen, and H. Li, “Learning structured sparsity in deep neural networks,” in
2016
Earlier work this paper cites.
A. See, M.-T. Luong, and C. D. Manning, “Compression of neural machine translation models via pruning,”
2016
Earlier work this paper cites.
S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,”
2016
Earlier work this paper cites.
D. Amodei, R. Anubhai, J. Bai, and E. Battenberg, et al., “Deep Speech 2: End-to-end speech recognition in English and Mandarin,” in
2016
Earlier work this paper cites.
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and Alexander C. Berg, “SSD: Single shot multibox detector,” in
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, and T. Rehfeld, et al., “The Cityscapes dataset for semantic urban scene understanding,” in
2016
Earlier work this paper cites.
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “SQuAD: 100,000+ questions for machine comprehension of text,” in
2016
Earlier work this paper cites.
X. Dong, J. Huang, Y. Yang, and S. Yan, “More is Less: A more complicated network with less inference complexity,” in
2017
Earlier work this paper cites.
J.-H. Luo, J. Wu, and W. Lin, “ThiNet: A filter level pruning method for deep neural network compression,” in
2017
Earlier work this paper cites.
Z. Liu, J. Li, Z. Shen, G. Huang, Shoumeng, and Z. Changshui, “Learning efficient convolutional networks through network slimming,” in
2017
Earlier work this paper cites.
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning filters for efficient convnets,” in
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, and N. Parmar, et al., “Attention is all you need,” in
2017
Earlier work this paper cites.
Y. He, X. Zhang, and J. Sun, “Channel pruning for accelerating very deep neural networks,” in
2017
Earlier work this paper cites.
P. Molchanov, S. Tyree, T. Karras, T. Aila, and J. Kautz, “Pruning convolutional neural networks for resource efficient inference,” in
2017
Earlier work this paper cites.
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in
2017
Earlier work this paper cites.
Z. Cao, T. Simon, S.-E. Wei, and Y. Sheikh, “Realtime multi-person 2D pose estimation using part affinity fields,” in
2017
Earlier work this paper cites.
K. Zhang, W. Zuo, Y. Chen, D. Meng, and L. Zhang, “Beyond a Gaussian denoiser: Residual learning of deep CNN for image denoising,”
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, and D. Summers-Stay, et al., “Making the V in VQA matter: Elevating the role of image understanding in visual question answering,” in
2017
Earlier work this paper cites.
S. Nrang, E. Elsen, G. Diamos, and S. Sengupta, “Exploring sparsity in recurrent neural networks,” in
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollar, “Focal loss for dense object detection,” in
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “GANs trained by a two time-scale update rule converge to a local nash equilibrium,” in
2017
Earlier work this paper cites.
J. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in
2017
Earlier work this paper cites.
I. Shankar, D. Nikhil, and C. Kornel, “First quora dataset release: Question pairs,” 2017. [Online]. Available:
2017
Earlier work this paper cites.
S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,” in
2017
Earlier work this paper cites.
S. Lin, R. Ji, C. Chen, D. Tao, and J. Luo, “Holistic CNN compression via low-rank decomposition with knowledge transfer,”
2018
Earlier work this paper cites.
Z. Huang and N. Wang, “Data-driven sparse structure selection for deep neural networks,” in
2018
Earlier work this paper cites.
A. Jacot, F. Gabriel, and C. Hongler, “Neural tangent kernel: Convergence and generalization in neural networks,” in
2018
Earlier work this paper cites.
Y. He, G. Kang, X. Dong, Y. Fu, and Y. Yang, “Soft filter pruning for accelerating deep convolutional neural networks,” in
2018
Earlier work this paper cites.
A. Gordon, E. Eban, O. Nachum, and B. Chen, “MorphNet:fast & simple resource-constrained structure learning of deep networks,” in
2018
Earlier work this paper cites.
D. C. Mocanu, E. Mocanu, P. Stone, P. H. Nguyen, M. Gibescu, and A. Liotta, “Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science,”
2018
Earlier work this paper cites.
X. Ding, G. Ding, J. Han, and S. Tang, “Auto-balanced filter pruning for efficient convolutional neural networks,” in
2018
Earlier work this paper cites.
Y. He, J. Lin, Z. Liu, H. Wang, L.-J. Li, and S. Han, “AMC: AutoML for model compression and acceleration on mobile devices,” in
2018
Earlier work this paper cites.
Z. Zhong, J. Yan, W. Wu, J. Shao, and C.-L. Liu, “Practical block-wise neural network architecture generation,” in
2018
Earlier work this paper cites.
T. Zhang, S. Ye, K. Zhang, J. Tang, W. Wen, M. Fardad, and Y. Wang, “A systematic DNN weight pruning framework using alternating direction method of multipliers,” in
2018
Earlier work this paper cites.
S. Kalyan, S. Joshi, S. Sheik, B. U. Pedroni, and G. Cauwcnbcrghs, “Unsupervised synaptic pruning strategies for restricted boltzmann machines,” in
2018
Earlier work this paper cites.
F. Tung and G. Mori, “CLIP-Q: Deep network compression learning by in-parallel pruning-quantization,” in
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Y. He, Y. Ding, P. Liu, L. Zhu, H. Zhang, and Y. Yang, “Learning filter pruning criteria for deep convolutional neural networks acceleration,” in
2018
Earlier work this paper cites.
E. Wong, F. R. Schmidt, J. Hendrik Metzen, and J. Zico Kolter, “Scaling provable adversarial defenses,” in
2018
Earlier work this paper cites.
A. Abdelhamed, S. Lin, and M. S. Brown, “A high-quality denoising dataset for smartphone cameras,” in
2018
Earlier work this paper cites.
P. Lison, J. Tiedemann, and M. Kouylekov, “OpenSubtitles2018: Statistical rescoring of sentence alignments in large, noisy parallel corpora,” in
2018
Earlier work this paper cites.
M. Zhu and S. Gupta, “To prune, or not to prune: exploring the efficacy of pruning for model compression,” in
2018
Earlier work this paper cites.
A. Williams, N. Nangia, and S. Bowman, “A broad-coverage challenge corpus for sentence understanding through inference,” in
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Earlier work this paper cites.
Z. You, K. Yan, J. Ye, M. Ma, and P. Wang, “Gate Decorator: Global filter pruning method for accelerating deep convolutional neural networks,” in
2019
Earlier work this paper cites.
H. Liu, K. Simonyan, and Y. Yang, “DARTS: Differentiable architecture search,” in
2019
Earlier work this paper cites.
J. Frankle and M. Carbin, “The lottery ticket hypothesis: finding sparse, trainable neural networks,” in
2019
Earlier work this paper cites.
N. Lee, T. Ajanthan, and P. H. Torr, “SNIP: Single-shot network pruning based on connection sensitivity,” in
2019
Earlier work this paper cites.
A. Gaier and D. Ha, “Weight agnostic neural networks,” in
2019
Earlier work this paper cites.
C. Zhao, B. Ni, J. Zhang, Q. Zhao, W. Zhang, and Q. Tian, “Variational convolutional neural network pruning,” in
2019
Earlier work this paper cites.
Z. Liu, H. Mu, X. Zhang, Z. Guo, X. Yang, T. K.-T. Cheng, and J. Sun, “MetaPruning: Meta learning for automatic neural network channel pruning,” in
2019
Earlier work this paper cites.
T. Li, B. Wu, Y. Yang, Y. Fan, Y. Zhang, and W. Liu, “Compressing convolutional neural networks via factorized convolutional filters,” in
2019
Earlier work this paper cites.
X. Dai, H. Yin, and N. K. Jha, “NeST: A neural network synthesis tool based on a grow-and-prune paradigm,”
2019
Earlier work this paper cites.
H. Mostafa and X. Wang, “Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization,” in
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. He, P. Liu, Z. Wang, Z. Hu, and Y. Yang, “Filter pruning via geometric median for deep convolutional neural networks acceleration,” in
2019
Earlier work this paper cites.
P. Molchanov, A. Mallya, S. Tyree, I. Frosio, and J. Kautz, “Importance estimation for neural network pruning,” in
2019
Earlier work this paper cites.
S. Lin, R. Ji, and Y. Li, et al., “Towards compact ConvNets via structure-sparsity regularized filter pruning,”
2019
Earlier work this paper cites.
Z. Liu, M. Sun, T. Zhou, G. Huang, and T. Darrell, “Rethinking the value of network pruning,” in
2019
Earlier work this paper cites.
A. S. Morcos, H. Yu, M. Paganini, and Y. Tian, “One ticket to win them all: Generalizing lottery ticket initializations across datasets and optimizers,” in
2019
Earlier work this paper cites.
R. Mehta, “Sparse transfer learning via winning lottery tickets,”
2019
Earlier work this paper cites.
S. Lin, R. Ji, C. Yan, B. Zhang, L. Cao, Q. Ye, F. Huang, and D. Doermann, “Towards optimal structured CNN pruning via generative adversarial learning,” in
2019
Earlier work this paper cites.
H. Yang, Y. Zhu, and J. Liu, “ECC: Platform-independent energy-constrained deep neural network compression via a bilinear regression model,” in
2019
Earlier work this paper cites.
P. Michel, O. Levy, and G. Neubig, “Are sixteen heads really better than one?” in
2019
Earlier work this paper cites.
Y. Rao, J. Lu, J. Lin, and J. Zhou, “Runtime network routing for efficient image classification,”
2019
Earlier work this paper cites.
W. Hua, Y. Zhou, C. D. Sa, Z. Zhang, and G. E. Suh, “Channel gating neural networks,” in
2019
Earlier work this paper cites.
X. Gao, Y. Zhao, and L. Dudziak, et al., “Dynamic channel pruning: Feature boosting and suppression,” in
2019
Earlier work this paper cites.
V. Sehwag, S. Wang, and P. Mittal, et al., “Towards compact and robust deep neural networks,”
2019
Earlier work this paper cites.
C. Wang, R. Grosse, and S. Fidler, et al., “EigenDamage: Structured pruning in the Kronecker-Factored eigenbasis,” in
2019
Earlier work this paper cites.
E. Voita, D. Talbot, F. Moiseev, R. Sennrich, and I. Titov, “Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned,” in
2019
Earlier work this paper cites.
Y. Li, S. Lin, B. Zhang, J. Liu, D. Doermann, Y. Wu, F. Huang, and R. Ji, “Exploiting kernel sparsity and entropy for interpretable CNN compression,” in
2019
Earlier work this paper cites.
X. Dong and Y. Yang, “Network pruning via transformable architecture search,” in
2019
Earlier work this paper cites.
C. Clark, K. Lee, M.-W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova, “BoolQ: Exploring the surprising difficulty of natural yes/no questions,” in
2019
Earlier work this paper cites.
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi, “Hellaswag: Can a machine really finish your sentence?” in
2019
Earlier work this paper cites.
X. Ding, G. Ding, X. Zhou, Y. Guo, J. Han, and J. Liu, “Global sparse momentum SGD for pruning very deep neural networks,” in
2019
Cited alongside, same era.
H. Shu, Y. Wang, X. Jia, K. Han, H. Chen, C. Xu, Q. Tian, and C. Xu, “Co-evolutionary compression for unpaired image translation,” in
2019
Cited alongside, same era.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “GLUE: A multi-task benchmark and analysis platform for natural language understanding,” in
2019
Cited alongside, same era.
H. You, C. Li, and P. Xu, et al., “Drawing early-bird tickets: Towards more efficient training of deep networks,” in
2020
Cited alongside, same era.
V. Sehwag, S. Wang, P. Mittal, and S. Jana, “HYDRA: Pruning adversarially robust neural networks,” in
2020
Cited alongside, same era.
Y. Bai, H. Wang, Z. Tao, K. Li, and Y. Fu, “Dual lottery ticket hypothesis,” in
2022
Later among the works it cites.
S. Liu, T. Chen, and Z. Atashgahi, et al., “Deep ensembling with no overhead for either training or testing: The all-round blessings of dynamic sparsity,” in
2022
Later among the works it cites.
G. Sokar, E. Mocanu, and D. C. Mocanu, et al., “Dynamic sparse training for deep reinforcement learning,” in
2022
Later among the works it cites.
U. Evci, Y. A. Ioannou, C. Keskin, and Y. Dauphin, “Gradient flow in sparse neural networks and how lottery tickets win,” in
2022
Later among the works it cites.
M. Ma, J. Wang, and Z. Yu, “Differentiable network pruning via polarization of probabilistic channelwise soft masks,”
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
D. Blalock, J. J. G. Ortiz, J. Frankle, and J. Guttag, “What is the state of neural network pruning?” in
2020
Cited alongside, same era.
J. T. O’ Neill, “An survey of neural network compression,”
2020
Cited alongside, same era.
2020
Cited alongside, same era.
T. Choudhary, V. Mishra, A. Goswami, and J. Sarangapani, “A comprehensive survey on model compression and acceleration,”
2020
Cited alongside, same era.
H. Tanaka, D. Kunin, D. L. Yamins, and S. Ganguli, “Pruning neural networks without any data by iteratively conserving synaptic flow,” in
2020
Cited alongside, same era.
F. Meng, H. Cheng, K. Li, H. Luo, X. Guo, G. Lu, and X. Sun, “Pruning filter in filter,” in
2020
Cited alongside, same era.
Z. Gan, Y.-C. Chen, L. Li, T. Chen, Y. Cheng, S. Wang, J. Liu, L. Wang, and Z. Liu, “Playing lottery tickets with vision and language,” in
2022
Later among the works it cites.
E. J. Hu, Y. Shen, and P. Wallis, et al., “LoRA: Low-rank adaptation of large language models,” in
2022
Later among the works it cites.
M. Xia, Z. Zhong, and D. Chen, “Structured pruning learns compact and accurate models,” in
2022
Later among the works it cites.
W. Kwon, S. kim, M. W. Mahoney, J. Hassoun, K. Keutzer, and A. Gholami, “A fast post-training pruning framework for transformers,” in
2022
Later among the works it cites.
S. Elkerdawy, M. Elhoushi, H. Zhang, and N. Ray, “Fire together wire together: A dynamic pruning approach with self-supervised mask prediction,” in
2022
Later among the works it cites.
J. Meng, L. Yang, J. Shin, D. Fan, and J. sun Seo, “Contrastive dual gating: Learning sparse features with contrastive learning,” in
2022
Later among the works it cites.
F. Yu, K. Huang, M. Wang, Y. Cheng, W. Chu, and L. Cui, “Width & depth pruning for vision transformers,” in
2022
Later among the works it cites.
E. Kurtic, D. Campos, and T. Nguyen, et al., “The optimal BERT surgeon: Scalable and accurate second-order pruning for large language models,” in
2022
Later among the works it cites.
M. Zhang, X. Yu, J. Rong, and L. Ou, “Graph pruning for model compression,”
2022
Later among the works it cites.
S. Zhang, S. Roller, and N. Goyal, et al., “OPT: Open pre-trained transformer language models,”
2022
Later among the works it cites.
C. R. Wolfe, Q. Wang, J. L. Kim, and A. Kyrillidis, “How much pre-training is enough to discover a good subnetwork?” in
2022
Later among the works it cites.
E. Iofinova, A. Peste, M. Kurtz, and D. Alistarh, “How well do sparse imagenet models transfer?” in
2022
Later among the works it cites.
C. Zhang, S. Bengio, and Y. Singer, “Are all layers created equal?”
2022
Later among the works it cites.
S. Pan, Y. Qin, T. Li, X. Li, and L. Hou, “Momentum contrastive pruning,” in
2022
Later among the works it cites.
J. Park and A. No, “Prune your model before distill it,” in
2022
Later among the works it cites.
L. Chen, Y. Chen, J. Xi, and X. Le, “Knowledge from the original network: restore a better pruned network with knowledge distillation,”
2022
Later among the works it cites.
Y. Zhang, Y. Yao, P. Ram, P. Zhao, T. Chen, M. Hong, Y. Wang, and S. Liu, “Advancing model pruning via bi-level optimization,” in
2022
Later among the works it cites.
W. Zou, Y. Wang, X. Fu, and Y. Cao, “Dreaming to prune image deraining networks,” in
2022
Later among the works it cites.
Q. Yan, D. Gong, Y. Liu, A. van den Hengel, and J. Q. Shi, “Learning bayesian sparse networks with full experience replay for continual learning,” in
2022
Later among the works it cites.
F. Corti, R. Entezari, S. Hooker, D. Bacciu, and O. Saukh, “Studying the impact of magnitude pruning on contrastive learning methods,” in
2022
Later among the works it cites.
Y. Jiang, S. Wang, V. Valls, B. J. Ko, W.-H. Lee, K. K. Leung, and L. Tassiulas, “Model pruning enables efficient federated learning on edge devices,”
2022
Later among the works it cites.
C. Zheng, zheyang li, and K. Zhang, et al., “SAViT: Structure-aware vision transformer pruning via collaborative optimization,” in
2022
Later among the works it cites.
L. Cai, Z. An, C. Yang, Y. Yan, and Y. Xu, “Prior gradient mask guided pruning-aware fine-tuning,” in
2022
Later among the works it cites.
Q. Zhang, S. Zuo, and C. Liang, et al., “PLATON: Pruning large transformer models with upper confidence bound of weight importance,” in
2022
Later among the works it cites.
M. Alizadeh, S. A. Tailor, and L. Zintgraf, et al., “Prospect pruning: Finding trainable weights at initialization using meta-gradients,” in
2022
Later among the works it cites.
Z. Hou and S.-Y. Kung, “Multi-dimensional model compression of vision transformer,” in
2022
Later among the works it cites.
S. Yu, T. Chen, and J. Shen, et al., “Unified visual transformer compression,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
R. Xu, F. Luo, C. Wang, B. Chang, J. Huang, S. Huang, and F. Huang, “From dense to sparse: Contrastive pruning for better pre-trained language model compression,” in
2022
Later among the works it cites.
M. Bonnaerens, M. Freiberger, and J. Dambre, “Anchor pruning for object detection,”
2022
Later among the works it cites.
W. Zhang, N. Wang, K. Chen, Y. Liu, and T. Zhao, “A pruning method for deep convolutional network based on heat map generation metrics,”
2022
Later among the works it cites.
Y. Li, M. Zhu, C. Luo, H. Weng, Y. Jang, T. Wei, and S.-T. Xia, “BAAT: Towards sample-specific backdoor attack with clean labels,” in
2022
Later among the works it cites.
S. Ding, T. Chen, and Z. Wang, “Audio lottery: Speech recognition made ultra-lightweight, transferable, and noise-robust,” in
2022
Later among the works it cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Latif, M. Shoukat, and F. Shamshad, et al., “Sparks of large audio models: A survey and outlook,”
2023
Closest in time.
J. Wu, W. Gan, Z. Chen, S. Wan, and H. Lin, “AI-generated content AIGC: A survey,”
2023
Closest in time.
2023
Closest in time.
X. Ma, G. Fang, and X. Wang, “LLM-Pruner: On the structural pruning of large language models,” in
2023
Closest in time.
2023
Closest in time.
Y. He and L. Xiao, “Structured pruning for deep convolutional neural networks: A survey,”
2023
Closest in time.
G. Fang, X. Ma, M. Song, M. B. Mi, and X. Wang, “DepGraph: Towards any structural pruning,” in
2023
Closest in time.
G. Fang, X. Ma, and X. Wang, “Structural pruning for diffusion models,” in
2023
Closest in time.
2023
Closest in time.
H. Wang and Y. Fu, “Trainability preserving neural structured pruning,” in
2023
Closest in time.
H. Yang, H. Yin, and M. Shen, et al., “Global vision transformer pruning with hessian-aware saliency,” in
2023
Closest in time.
D. Hoang, S. Liu, and R. Marculescu, et al., “Revisiting pruning at initialization through the lens of Ramanujan graph,” in
2023
Closest in time.
M. Cho, S. Adya, and D. Naik, “PDP: Parameter-free differentiable pruning is all you need,” in
2023
Closest in time.
L. Yu and W. Xiang, “X-pruner: explainable pruning for vision transformers,” in
2023
Closest in time.
B. Hui, D. Yan, X. Ma, and W.-S. Ku, “Rethinking graph lottery tickets: Graph sparsity matters,” in
2023
Closest in time.
M. Paul, F. Chen, and B. W. Larsen, et al., “Unmasking the lottery ticket hypothesis-what’s encoded in a winning ticket’s mask,” in
2023
Closest in time.
D. Shi, C. Tao, Y. Jin, Z. Yang, C. Yuan, and J. Wang, “UPop: Unified and progressive pruning for compressing vision-language transformers,” in
2023
Closest in time.
A. Nova, H. Dai, and D. Schuurmans, “Gradient-free structured pruning with unlabeled data,” in
2023
Closest in time.
S. Tuli and N. K. Jha, “AccelTran: A sparsity-aware accelerator for dynamic inference with transformers,”
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Y. Li, Y. Yu, and Q. Zhang, et al., “LoSparse: Structured compression of large language models based on low-rank and sparse approximation,” in
2023
Closest in time.
A. Klein, J. Golebiowski, X. Ma, V. Perrone, and C. Archambeau, “Structural pruning of large language models via neural architecture search,” in
2023
Closest in time.
2023
Closest in time.
X. Sui, Q. Lv, L. Zhi, B. Zhu, Y. Yang, Y. Zhang, and Z. Tan, “A hardware-friendly high-precision cnn pruning method and its fpga implementation,”
2023
Closest in time.
OpenAI, “Gpt-4 technical report,” OpenAI, Tech. Rep., 2023
2023
Closest in time.
2023
Closest in time.
Y. Zhang, M. Lin, and Y. Zhong, et al., “Lottery jackpots exist in pre-trained models,”
2023
Closest in time.
Y. Wang, D. Li, and R. Sun, “NTK-SAP: Improving neural network pruning by aligning training dynamics,” in
2023
Closest in time.
T. Narshana, C. Murti, and C. Bhattacharyya, “DFPC: Data flow driven pruning of coupled channels without data,” in
2023
Closest in time.
M. Yin, B. Uzkent, and Y. Shen, et al., “GOHSP: A unified framework of graph and optimization-based heterogeneous structured pruning for vision transformer,” in
2023
Closest in time.
A. Jaiswal, S. Liu, and T. Chen, et al., “Instant soup: Cheap pruning ensembles in a single pass can draw lottery tickets from large models,” in
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
X. Shi, X. Ning, and L. Guo, et al., “Memory-oriented structural pruning for efficient image restoration,” in
2023
Closest in time.
H. Jiang, L. L. Zhang, and Y. Li, et al., “Accurate and structured pruning for efficient automatic speech recognition,” in
2023
Closest in time.
J. Lin, H. Yin, and W. Ping, et al., “Vila: On pre-training for visual language models,” in
2024
Closest in time.
2024
Closest in time.
S. Ashkboos, M. L. Croci, and M. Gennari, et al., “SliceGPT: Compress large language models by deleting rows and columns,” in
2024
Closest in time.
W. Shao, M. Chen, and Z. Zhang, et al., “OmniQuant: Omnidirectionally calibrated quantization for large language models,” in
2024
Closest in time.
Y. Gu, L. Dong, F. Wei, and M. Huang, “MiniLLM: Knowledge distillation of large language models,” in
2024
Closest in time.
X. Xu, M. Li, and C. Tao, et al., “A survey on knowledge distillation of large language models,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
M. Sun, Z. Liu, A. Bair, and J. Z. Kolter, “A simple and effective pruning approach for large language models,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
B. Liu, Z. Zhang, and P. He, et al., “A survey of lottery ticket hypothesis,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
M. Xia, T. Gao, Z. Zeng, and D. Chen, “Sheared LLaMA: Accelerating language model pre-training via structured pruning,” in
2024
Closest in time.
Y. An, X. Zhao, and T. Yu, et al., “Fluctuation-based adaptive structured pruning for large language models,” in
2024
Closest in time.
T. F. A. van der Ouderaa, M. Nagel, and M. van Baalen, et al., “The LLM surgeon,”
2024
Closest in time.
I. Amos, J. Berant, and A. Gupta, “Never train from scratch: Fair comparison of long-sequence models requires data-driven priors,” in
2024
Closest in time.
2024
Closest in time.
R. Gong, Y. Yong, and Z. Wang, et al., “Fast and controllable post-training sparsity-learning optimal sparsity allocation with global constraint in minutes,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Yang, Z. Cao, and H. Zhao, “LaCo: Large language model pruning via layer collapse,”
2024
Closest in time.
K. Han, Y. Wang, J. Guo, and E. Wu, “ParameterNet: Parameters are all you need,” in
2024
Closest in time.
H. He, J. Cai, and J. Liu, et al., “Pruning self-attentions into convolutional layers in single path,”
2024
Closest in time.
X. Wu, S. Gao, and Z. Zhang, et al., “Auto-train-once: Controller network guided automatic network pruning from scratch,” in
2024
Closest in time.
Y. He and J. T. Zhou, “Data-independent module-aware pruning for hierarchical vision transformers,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y.-L. Sung, J. Yoon, and M. Bansal, “ECoFLaP: Efficient coarse-to-fine layer-wise pruning for vision-language models,” in
2024
Closest in time.