Fetching the paper…
Reading the bibliography…
Neural Networks can be effectively compressed through pruning, significantly reducing storage and compute demands while maintaining predictive performance.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1989
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B. and Stork, D · 1992
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Earlier work this paper cites.
To prune, or not to prune: Exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A · 2018
Earlier work this paper cites.
K for the price of 1: Parameter-efficient multi-task and transfer learning
Mudrakarta, P. K., Sandler, M., Zhmoginov, A., and Howard, A · 2018
Earlier work this paper cites.
Sparse networks from scratch: Faster training without losing performance
Dettmers, T. and Zettlemoyer, L · 2019
Earlier work this paper cites.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Pruning by explaining: A novel criterion for deep neural network pruning
Yeom, S.-K., Seegerer, P., Lapuschkin, S., Binder, A., Wiedemann, S., Müller, K.-R., and Samek, W · 2019
Earlier work this paper cites.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Aghajanyan, A., Zettlemoyer, L., and Gupta, S · 2020
Earlier work this paper cites.
Rigging the lottery: Making all tickets winners
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Earlier work this paper cites.
Training batchnorm and only batchnorm: On the expressive power of random features in cnns
Frankle, J., Schwab, D. J., and Morcos, A. S · 2020
Cited alongside, same era.
Layer-adaptive sparsity for the magnitude-based pruning
Lee, J., Park, S., Mo, S., Ahn, S., and Shin, J · 2020
Cited alongside, same era.
Eagleeye: Fast sub-net evaluation for efficient neural network pruning
Li, B., Wu, B., Su, J., and Wang, G · 2020
Cited alongside, same era.
Dynamic model pruning with feedback
Lin, T., Stich, S. U., Barba, L., Dmitriev, D., and Jaggi, M · 2020
Cited alongside, same era.
Deep neural network training with frank-wolfe
Pokutta, S., Spiegel, C., and Zimmer, M · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
A fast post-training pruning framework for transformers
Kwon, W., Kim, S., Mahoney, M. W., Hassoun, J., Keutzer, K., and Gholami, A · 2022
Later among the works it cites.
Parameter-efficient sparsity for large language models fine-tuning
Li, Y., Luo, F., Tan, C., Wang, M., Huang, S., Li, S., and Bai, J · 2022
Later among the works it cites.
Efficient fine-tuning of bert models on the edge
Vucetic, D., Tayaranian, M., Ziaeefard, M., Clark, J. J., Meyer, B. H., and Gross, W. J · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2022
Later among the works it cites.
Compression-aware training of neural networks using frank-wolfe
Zimmer, M., Spiegel, C., and Pokutta, S · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Comparing rewinding and fine-tuning in neural network pruning
Renda, A., Frankle, J., and Carbin, M · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A · 2020
Cited alongside, same era.
Drawing early-bird tickets: Toward more efficient training of deep networks
You, H., Li, C., Xu, P., Fu, Y., Wang, Y., Chen, X., Baraniuk, R. G., Wang, Z., and Lin, Y · 2020
Cited alongside, same era.
Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Network pruning that matters: A case study on retraining variants
Le, D. H. and Hua, B.-S · 2021
Cited alongside, same era.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Closest in time.
A framework for few-shot language model evaluation, 12 2023
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2023
Closest in time.
The expressive power of tuning only the normalization layers
Giannou, A., Rajput, S., and Papailiopoulos, D · 2023
Closest in time.
Mistral 7b
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Closest in time.
Llm-pruner: On the structural pruning of large language models
Ma, X., Fang, G., and Wang, X · 2023
Closest in time.
A simple and effective pruning approach for large language models
Sun, M., Liu, Z., Bair, A., and Kolter, J. Z · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Closest in time.
How does calibration data affect the post-training pruning and quantization of large language models?
Williams, M. and Aletras, N · 2023
Closest in time.
Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity
Yin, L., Wu, Y., Zhang, Z., Hsieh, C.-Y., Wang, Y., Jia, Y., Pechenizkiy, M., Liang, Y., Wang, Z., and Liu, S · 2023
Closest in time.
How I Learned To Stop Worrying And Love Retraining
Zimmer, M., Spiegel, C., and Pokutta, S · 2023
Closest in time.
Spp: Sparsity-preserved parameter-efficient fine-tuning for large language models
Lu, X., Zhou, A., Xu, Y., Zhang, R., Gao, P., and Li, H · 2024
Closest in time.
Mixtral of experts — mistral.ai
Mistral, M. A · 2024
Closest in time.
Sqft: Low-cost model adaptation in low-precision sparse foundation models
Muñoz, J. P., Yuan, J., and Jain, N · 2024
Closest in time.
Towards meta-pruning via optimal transport
Theus, A., Geimer, O., Wicke, F., Hofmann, T., Anagnostidis, S., and Singh, S. P · 2024
Closest in time.
Sparse model soups: A recipe for improved pruning via model averaging
Zimmer, M., Spiegel, C., and Pokutta, S · 2024
Closest in time.