Fetching the paper…
Reading the bibliography…
Sparse activation, which selectively activates only an input-dependent set of neurons in inference, is a useful technique to reduce the computing cost of Large Language Models (LLMs) without retraining or adaptation efforts.
Axiomatic attribution for deep networks
M. Sundararajan, A. Taly, and Q. Yan · 2017
Earlier work this paper cites.
Snip: Single-shot network pruning based on connection sensitivity
N. Lee, T. Ajanthan, and P. H. Torr · 2018
Earlier work this paper cites.
Are sixteen heads really better than one?
P. Michel, O. Levy, and G. Neubig · 2019
Earlier work this paper cites.
Inducing and exploiting activation sparsity for fast inference on deep neural networks
M. Kurtz, J. Kopinsky, R. Gelashvili, A. Matveev, J. Carr, M. Goin, W. Leiserson, S. Moore, N. Shavit, and D. Alistarh · 2020
Earlier work this paper cites.
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks
T. Hoefler, D. Alistarh, T. Ben-Nun, N. Dryden, and A. Peste · 2021
Earlier work this paper cites.
Perturbation-based methods for explaining deep neural networks: A survey
M. Ivanovs, R. Kadikis, and K. Ozols · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2021
Earlier work this paper cites.
Group fisher pruning for practical network compression
L. Liu, S. Zhang, Z. Kuang, A. Zhou, J.-H. Xue, X. Wang, Y. Chen, W. Yang, Q. Liao, and W. Zhang · 2021
Earlier work this paper cites.
Moefication: Transformer feed-forward layers are mixtures of experts
Z. Zhang, Y. Lin, Z. Liu, P. Li, M. Sun, and J. Zhou · 2021
Earlier work this paper cites.
Attribution-based xai methods in computer vision: A review
K. Abhishek and D. Kamath · 2022
Earlier work this paper cites.
H. Bansal, K. Gopalakrishnan, S. Dingliwal, S. Bodapati, K. Kirchhoff, and D. Roth · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Cited alongside, same era.
Singe: Sparsity via integrated gradients estimation of neuron relevance
E. Yvinec, A. Dapogny, M. Cord, and K. Bailly · 2022
Cited alongside, same era.
The falcon series of open language models
E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojocaru, M. Debbah, É. Goffinet, D. Hesslow, J. Launay, Q. Malartic, et al · 2023
Cited alongside, same era.
S. Gunasekar, Y. Zhang, J. Aneja, C. C. T. Mendes, A. Del Giorno, S. Gopi, M. Javaheripi, P. Kauffmann, G. de Rosa, O. Saarikivi, et al · 2023
Cited alongside, same era.
Simplifying transformer blocks
B. He and T. Hofmann · 2023
Powerinfer: Fast large language model serving with a consumer-grade gpu
Y. Song, Z. Mi, H. Xie, and H. Chen · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Multistage collaborative knowledge distillation from large language models
J. Zhao, W. Zhao, A. Drozdov, B. Rozonoyer, M. A. Sultan, J.-Y. Lee, M. Iyyer, and A. McCallum · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. Behl, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
P. Ke, B. Wen, Z. Feng, X. Liu, X. Lei, J. Cheng, S. Wang, A. Zeng, Y. Dong, H. Wang, et al · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Y. Leviathan, M. Kalman, and Y. Matias · 2023
Cited alongside, same era.
Awq: Activation-aware weight quantization for llm compression and acceleration
J. Lin, J. Tang, H. Tang, S. Yang, X. Dang, and S. Han · 2023
Cited alongside, same era.
Deja vu: Contextual sparsity for efficient llms at inference time
Z. Liu, J. Wang, T. Dao, T. Zhou, B. Yuan, Z. Song, A. Shrivastava, C. Zhang, Y. Tian, C. Re, et al · 2023
Cited alongside, same era.
Llm-pruner: On the structural pruning of large language models
X. Ma, G. Fang, and X. Wang · 2023
Cited alongside, same era.
Phi-2: The surprising power of small language models
S. B. M. G. Mojan Javaheripi · 2023
Cited alongside, same era.
https://huggingface.co/datasets/yahoo_answers_qa
Yahoo answers qa dataset
Cited in the paper.
Quip: 2-bit quantization of large language models with guarantees
J. Chee, Y. Cai, V. Kuleshov, and C. M. De Sa · 2024
Closest in time.
Knowledge-augmented reasoning distillation for small language models in knowledge-intensive tasks
M. Kang, S. Lee, J. Baek, K. Kawaguchi, and S. J. Hwang · 2024
Closest in time.
Memory-efficient fine-tuning of compressed large language models via sub-4-bit integer quantization
J. Kim, J. H. Lee, S. Kim, J. Park, K. M. Yoo, S. J. Kwon, and D. Lee · 2024
Closest in time.
Ziplm: Inference-aware structured pruning of language models
E. Kurtić, E. Frantar, and D. Alistarh · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al · 2024
Closest in time.
Mobillama: Towards accurate and lightweight fully transparent gpt
O. Thawakar, A. Vayani, S. Khan, H. Cholakal, R. M. Anwer, M. Felsberg, T. Baldwin, E. P. Xing, and F. S. Khan · 2024
Closest in time.
Relu 2 wins: Discovering efficient activation functions for sparse llms
Z. Zhang, Y. Song, G. Yu, X. Han, Y. Lin, C. Xiao, C. Song, Z. Liu, Z. Mi, and M. Sun · 2024
Closest in time.