Fetching the paper…
Reading the bibliography…
Dynamic activation (DA) techniques, such as DejaVu and MoEfication, have demonstrated their potential to significantly enhance the inference efficiency of large language models (LLMs).
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Clark, C.; Lee, K.; Chang, M.-W.; Kwiatkowski, T.; Collins, M.; and Toutanova, K. 2019 · 1905
Earlier work this paper cites.
HellaSwag: Can a Machine Really Finish Your Sentence?
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 1905
Earlier work this paper cites.
PIQA: Reasoning about Physical Commonsense in Natural Language
Bisk, Y.; Zellers, R.; Bras, R. L.; Gao, J.; and Choi, Y. 2019 · 1911
Earlier work this paper cites.
Proving the Lottery Ticket Hypothesis: Pruning is All You Need
Malach, E.; Yehudai, G.; Shalev-Shwartz, S.; and Shamir, O. 2020 · 2002
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Roemmele, M.; Bejan, C. A.; and Gordon, A. S. 2011 · 2011
Earlier work this paper cites.
Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond
Nallapati, R.; Zhou, B.; dos santos, C. N.; Gulcehre, C.; and Xiang, B. 2016 · 2016
Earlier work this paper cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Shazeer, N.; Mirhoseini, A.; Maziarz, K.; Davis, A.; Le, Q.; Hinton, G.; and Dean, J. 2017 · 2017
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Earlier work this paper cites.
Narayan, S.; Cohen, S. B.; and Lapata, M. 2018 · 2018
Earlier work this paper cites.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Frankle, J.; and Carbin, M. 2019 · 2019
Earlier work this paper cites.
Accelerating Convolutional Neural Networks via Activation Map Compression
Georgiadis, G. 2019 · 2019
Earlier work this paper cites.
CoQA: A Conversational Question Answering Challenge
Reddy, S.; Chen, D.; and Manning, C. D. 2019 · 2019
Earlier work this paper cites.
Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural Networks
Kurtz, M.; Kopinsky, J.; Gelashvili, R.; Matveev, A.; Carr, J.; Goin, M.; Leiserson, W.; Moore, S.; Shavit, N.; and Alistarh, D. 2020 · 2020
Cited alongside, same era.
On the Opportunities and Risks of Foundation Models
Bommasani, R.; Hudson, D. A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M. S.; Bohg, J.; Bosselut, A.; Brunskill, E.; Brynjolfsson, E.; Buch, S.; Card, D.; Castellon, R.; Chatterji, N.; Chen, A.; Creel, K.; Davis, J. Q.; Demszky, D.; Donahue, C.; Doumbouya, M.; Durmus, E.; Ermon, S.; Etchemendy, J.; Ethayarajh, K.; Fei-Fei, L.; Finn, C.; Gale, T.; Gillespie, L.; Goel, K.; Goodman, N.; Grossman, S.; Guha, N.; Hashimoto, T.; Henderson, P.; Hewitt, J.; Ho, D. E.; Hong, J.; Hsu, K.; Huang, J.; Icard, T.; Jain, S.; Jurafsky, D.; Kalluri, P.; Karamcheti, S.; Keeling, G.; Khani, F.; Khattab, O.; Koh, P. W.; Krass, M.; Krishna, R.; Kuditipudi, R.; Kumar, A.; Ladhak, F.; Lee, M.; Lee, T.; Leskovec, J.; Levent, I.; Li, X. L.; Li, X.; Ma, T.; Malik, A.; Manning, C. D.; Mirchandani, S.; Mitchell, E.; Munyikwa, Z.; Nair, S.; Narayan, A.; Narayanan, D.; Newman, B.; Nie, A.; Niebles, J. C.; Nilforoshan, H.; Nyarko, J.; Ogut, G.; Orr, L.; Papadimitriou, I.; Park, J. S.; Piech, C.; Portelance, E.; Potts, C.; Raghunathan, A.; Reich, R.; Ren, H.; Rong, F.; Roohani, Y.; Ruiz, C.; Ryan, J.; Ré, C.; Sadigh, D.; Sagawa, S.; Santhanam, K.; Shih, A.; Srinivasan, K.; Tamkin, A.; Taori, R.; Thomas, A. W.; Tramèr, F.; Wang, R. E.; Wang, W.; Wu, B.; Wu, J.; Wu, Y.; Xie, S. M.; Yasunaga, M.; You, J.; Zaharia, M.; Zhang, M.; Zhang, T.; Zhang, X.; Zhang, Y.; Zheng, L.; Zhou, K.; and Liang, P. 2022 · 2022
Cited alongside, same era.
SCROLLS: Standardized CompaRison Over Long Language Sequences
Shaham, U.; Segal, E.; Ivgi, M.; Efrat, A.; Yoran, O.; Haviv, A.; Gupta, A.; Xiong, W.; Geva, M.; Berant, J.; and Levy, O. 2022 · 2022
Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
Li, P.; Zhang, Z.; Yadav, P.; Sung, Y.-L.; Cheng, Y.; Bansal, M.; and Chen, T. 2024 · 2024
Closest in time.
Dynamic Activation Pitfalls in LLaMA Models: An Empirical Study
Ma, C.; Huang, M.; Wang, C.; Wang, Y.; and Yu, L. 2024 · 2024
Closest in time.
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
Pan, B.; Shen, Y.; Liu, H.; Mishra, M.; Zhang, G.; Oliva, A.; Raffel, C.; and Panda, R. 2024 · 2024
Closest in time.
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
Song, C.; Han, X.; Zhang, Z.; Hu, S.; Shi, X.; Li, K.; Chen, C.; Liu, Z.; Li, G.; Yang, T.; and Sun, M. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Frantar, E.; and Alistarh, D. 2023 · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation
Gao, L.; Tow, J.; Abbasi, B.; Biderman, S.; Black, S.; DiPofi, A.; Foster, C.; Golding, L.; Hsu, J.; Le Noac’h, A.; Li, H.; McDonell, K.; Muennighoff, N.; Ociepa, C.; Phang, J.; Reynolds, L.; Schoelkopf, H.; Skowron, A.; Sutawika, L.; Tang, E.; Thite, A.; Wang, B.; Wang, K.; and Zou, A. 2023 · 2023
Cited alongside, same era.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; de las Casas, D.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; Lavaud, L. R.; Lachaux, M.-A.; Stock, P.; Scao, T. L.; Lavril, T.; Wang, T.; Lacroix, T.; and Sayed, W. E. 2023 · 2023
Cited alongside, same era.
The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers
Li, Z.; You, C.; Bhojanapalli, S.; Li, D.; Rawat, A. S.; Reddi, S. J.; Ye, K.; Chern, F.; Yu, F.; Guo, R.; and Kumar, S. 2023 · 2023
Cited alongside, same era.
ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models
Mirzadeh, I.; Alizadeh, K.; Mehta, S.; Mundo, C. C. D.; Tuzel, O.; Samei, G.; Rastegari, M.; and Farajtabar, M. 2023 · 2023
Cited alongside, same era.
H 2 O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
Zhang, Z.; Sheng, Y.; Zhou, T.; Chen, T.; Zheng, L.; Cai, R.; Song, Z.; Tian, Y.; Ré, C.; Barrett, C.; Wang, Z.; and Chen, B. 2023 · 2023
Cited alongside, same era.
STAR: Sparse Thresholded Activation under partial-Regularization for Activation Sparsity Exploration
Zhu, Z.; Pourtaherian, A.; Waeijen, L.; Bondarev, E.; and Moreira, O. 2023 · 2023
Cited alongside, same era.
SliceGPT: Compress Large Language Models by Deleting Rows and Columns
Ashkboos, S.; Croci, M. L.; do Nascimento, M. G.; Hoefler, T.; and Hensman, J. 2024 · 2024
Cited alongside, same era.
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
Dong, H.; Chen, B.; and Chi, Y. 2024 · 2024
Cited alongside, same era.
Szatkowski, F.; Wójcik, B.; Piórczyński, M.; and Scardapane, S. 2024 · 2024
Closest in time.
Gemma: Open Models Based on Gemini Research and Technology
Team, G.; Mesnard, T.; Hardin, C.; Dadashi, R.; Bhupatiraju, S.; Pathak, S.; Sifre, L.; Rivière, M.; Kale, M. S.; Love, J.; Tafti, P.; Hussenot, L.; Sessa, P. G.; Chowdhery, A.; Roberts, A.; Barua, A.; Botev, A.; Castro-Ros, A.; Slone, A.; Héliou, A.; Tacchetti, A.; Bulanova, A.; Paterson, A.; Tsai, B.; Shahriari, B.; Lan, C. L.; Choquette-Choo, C. A.; Crepy, C.; Cer, D.; Ippolito, D.; Reid, D.; Buchatskaya, E.; Ni, E.; Noland, E.; Yan, G.; Tucker, G.; Muraru, G.-C.; Rozhdestvenskiy, G.; Michalewski, H.; Tenney, I.; Grishchenko, I.; Austin, J.; Keeling, J.; Labanowski, J.; Lespiau, J.-B.; Stanway, J.; Brennan, J.; Chen, J.; Ferret, J.; Chiu, J.; Mao-Jones, J.; Lee, K.; Yu, K.; Millican, K.; Sjoesund, L. L.; Lee, L.; Dixon, L.; Reid, M.; Mikuła, M.; Wirth, M.; Sharman, M.; Chinaev, N.; Thain, N.; Bachem, O.; Chang, O.; Wahltinez, O.; Bailey, P.; Michel, P.; Yotov, P.; Chaabouni, R.; Comanescu, R.; Jana, R.; Anil, R.; McIlroy, R.; Liu, R.; Mullins, R.; Smith, S. L.; Borgeaud, S.; Girgin, S.; Douglas, S.; Pandya, S.; Shakeri, S.; De, S.; Klimenko, T.; Hennigan, T.; Feinberg, V.; Stokowiec, W.; hui Chen, Y.; Ahmed, Z.; Gong, Z.; Warkentin, T.; Peran, L.; Giang, M.; Farabet, C.; Vinyals, O.; Dean, J.; Kavukcuoglu, K.; Hassabis, D.; Ghahramani, Z.; Eck, D.; Barral, J.; Pereira, F.; Collins, E.; Joulin, A.; Fiedel, N.; Senter, E.; Andreev, A.; and Kenealy, K. 2024 · 2024
Closest in time.
LLM Inference Unveiled: Survey and Roofline Model Insights
Yuan, Z.; Shang, Y.; Zhou, Y.; Dong, Z.; Zhou, Z.; Xue, C.; Wu, B.; Li, Z.; Gu, Q.; Lee, Y. J.; Yan, Y.; Chen, B.; Sun, G.; and Keutzer, K. 2024 · 2024
Closest in time.
ReLU 2 Wins: Discovering Efficient Activation Functions for Sparse LLMs
Zhang, Z.; Song, Y.; Yu, G.; Han, X.; Lin, Y.; Xiao, C.; Song, C.; Liu, Z.; Mi, Z.; and Sun, M. 2024 · 2024
Closest in time.
Learn To be Efficient: Build Structured Sparsity in Large Language Models
Zheng, H.; Bai, X.; Liu, X.; Mao, Z. M.; Chen, B.; Lai, F.; and Prakash, A. 2024 · 2024
Closest in time.
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
Zhong, Z.; Xia, M.; Chen, D.; and Lewis, M. 2024 · 2024
Closest in time.
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
Zhu, T.; Qu, X.; Dong, D.; Ruan, J.; Tong, J.; He, C.; and Cheng, Y. 2024 · 2024
Closest in time.