Fetching the paper…
Reading the bibliography…
We present Junk DNA Hypothesis by adopting a novel task-centric angle for the pre-trained weights of large language models (LLMs).
So much "junk" dna in our genome
Ohno, S · 1972
Earlier work this paper cites.
Using relevance to reduce network size automatically
Mozer, M. C. and Smolensky, P · 1989
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B. and Stork, D · 1992
Earlier work this paper cites.
Role of conserved non-coding dna elements in the foxp3 gene in regulatory t-cell fate
Zheng, Y., Josefowicz, S., Chaudhry, A., Peng, X. P., Forbush, K., and Rudensky, A. Y · 2010
Earlier work this paper cites.
Junk DNA: a journey through the dark matter of the genome
Carey, N · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2016
Earlier work this paper cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Dong, X., Chen, S., and Pan, S · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L · 2017
Earlier work this paper cites.
On compressing deep models by low rank and sparse decomposition
Yu, X., Liu, T., Wang, X., and Tao, D · 2017
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
Mlprune: Multi-layer pruning for automated neural network compression
Zeng, W. and Urtasun, R · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2019
Earlier work this paper cites.
The lottery ticket hypothesis at scale
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M · 2019
Earlier work this paper cites.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S · 2019
Earlier work this paper cites.
FreebaseQA: A new factoid QA data set matching trivia-style question-answer pairs with Freebase
Jiang, K., Wu, D., and Jiang, H · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Mostafa, H. and Wang, X · 2019
Earlier work this paper cites.
Human vs. muppet: A conservative estimate of human performance on the glue benchmark
Nangia, N. and Bowman, S. R · 2019
Cited alongside, same era.
Eigendamage: Structured pruning in the kronecker-factored eigenbasis
Wang, C., Grosse, R., Fidler, S., and Zhang, G · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Do we actually need dense over-parameterization? in-time over-parameterization in sparse training
Liu, S., Yin, L., Mocanu, D. C., and Pechenizkiy, M · 2021
Later among the works it cites.
Rethinking network pruning–under the pre-train and fine-tune paradigm
Xu, D., Yen, I. E., Zhao, J., and Xiao, Z · 2021
Later among the works it cites.
Mest: Accurate and fast memory-economic sparse training framework on the edge
Yuan, G., Ma, X., Niu, W., Li, Z., Kong, Z., Liu, N., Gong, Y., Zhan, Z., He, C., Jin, Q., et al · 2021
Later among the works it cites.
Prune once for all: Sparse pre-trained language models
Zafrir, O., Larey, A., Boudoukh, G., Shen, H., and Wasserblat, M · 2021
Later among the works it cites.
Learning n: m fine-grained structured sparse neural networks from scratch
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The lottery ticket hypothesis for pre-trained bert networks
Chen, T., Frankle, J., Chang, S., Liu, S., Zhang, Y., Wang, Z., and Carbin, M · 2020
Cited alongside, same era.
Rigging the lottery: Making all tickets winners
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Cited alongside, same era.
Linear mode connectivity and the lottery ticket hypothesis
Frankle, J., Dziugaite, G. K., Roy, D., and Carbin, M · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Cited alongside, same era.
Nvidia a100 tensor core gpu architecture
Nvidia · 2020
Cited alongside, same era.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V., Wolf, T., and Rush, A · 2020
Cited alongside, same era.
Woodfisher: Efficient second-order approximation for neural network compression
Singh, S. P. and Alistarh, D · 2020
Cited alongside, same era.
Zhou, A., Ma, Y., Zhu, J., Liu, J., Zhang, Z., Yuan, K., Sun, W., and Li, H · 2021
Later among the works it cites.
Training your sparse neural network better with any mask
Jaiswal, A. K., Ma, H., Chen, T., Ding, Y., and Wang, Z · 2022
Later among the works it cites.
The optimal bert surgeon: Scalable and accurate second-order pruning for large language models
Kurtic, E., Campos, D., Nguyen, T., Frantar, E., Kurtz, M., Fineran, B., Goin, M., and Alistarh, D · 2022
Later among the works it cites.
Large models are parsimonious learners: Activation sparsity in trained transformers
Li, Z., You, C., Bhojanapalli, S., Li, D., Rawat, A. S., Reddi, S. J., Ye, K., Chern, F., Yu, F., Guo, R., et al · 2022
Later among the works it cites.
Liu, S., Chen, T., Chen, X., Shen, L., Mocanu, D. C., Wang, Z., and Pechenizkiy, M · 2022
Later among the works it cites.
Platon: Pruning large transformer models with upper confidence bound of weight importance
Zhang, Q., Zuo, S., Liang, C., Bukharin, A., He, P., Chen, W., and Zhao, T · 2022
Later among the works it cites.
Spqr: A sparse-quantized representation for near-lossless llm weight compression
Dettmers, T., Svirschevski, R., Egiazarian, V., Kuznedelev, D., Frantar, E., Ashkboos, S., Borzunov, A., Hoefler, T., and Alistarh, D · 2023
Closest in time.
Sparsegpt: Massive language models can be accurately pruned in one-shot, 2023
Frantar, E. and Alistarh, D · 2023
Closest in time.
Sparsity may cry: Let us fail (current) sparse neural networks together!
Liu, S., Chen, T., Zhang, Z., Chen, X., Huang, T., Jaiswal, A., and Wang, Z · 2023
Closest in time.
Lightformer: Light-weight transformer using svd-based weight transfer and parameter sharing
Lv, X., Zhang, P., Li, S., Gan, G., and Sun, Y · 2023
Closest in time.
Llm-pruner: On the structural pruning of large language models
Ma, X., Fang, G., and Wang, X · 2023
Closest in time.
Few-shot fine-tuning vs. in-context learning: A fair comparison and evaluation
Mosbach, M., Pimentel, T., Ravfogel, S., Klakow, D., and Elazar, Y · 2023
Closest in time.
In-context retrieval-augmented language models
Ram, O., Levine, Y., Dalmedigos, I., Muhlgay, D., Shashua, A., Leyton-Brown, K., and Shoham, Y · 2023
Closest in time.
The truth is in there: Improving reasoning in language models with layer-selective rank reduction
Sharma, P., Ash, J. T., and Misra, D · 2023
Closest in time.
A simple and effective pruning approach for large language models
Sun, M., Liu, Z., Bair, A., and Kolter, J. Z · 2023
Closest in time.
Dynamic sparsity is channel-level sparsity learner
Yin, L., Li, G., Fang, M., Shen, L., Huang, T., Wang, Z., Menkovski, V., Ma, X., Pechenizkiy, M., and Liu, S · 2023
Closest in time.