Fetching the paper…
Reading the bibliography…
Sparse computation offers a compelling solution for the inference of Large Language Models (LLMs) in low-resource scenarios by dynamically skipping the computation of inactive neurons.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E · 2010
Earlier work this paper cites.
Deep learning of representations: Looking forward
Bengio, Y · 2013
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Dynamic models of large-scale brain activity
Breakspear, M · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., and Uszkoreit, J · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the AI2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Elfwing, S., Uchibe, E., and Doya, K · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? A new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2019
Earlier work this paper cites.
PIQA: reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2020
Earlier work this paper cites.
SLIDE : In defense of smart algorithms over hardware acceleration for large-scale deep learning systems
Chen, B., Medini, T., Farwell, J., Gobriel, S., Tai, T. C., and Shrivastava, A · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified Text-to-Text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2020
Earlier work this paper cites.
GLU variants improve transformer
Shazeer, N · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Language models are Few-Shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2021
Earlier work this paper cites.
Scatterbrain: Unifying sparse and low-rank attention
Chen, B., Dao, T., Winsor, E., Song, Z., Rudra, A., and Ré, C · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Turbotransformers: an efficient GPU serving system for transformer models
Fang, J., Yu, Y., Zhao, C., and Zhou, J · 2021
Earlier work this paper cites.
A framework for few-shot language model evaluation, September 2021
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2021
Cited alongside, same era.
Transmission delays and frequency detuning can regulate information flow between brain regions
Pariz, A., Fischer, I., Valizadeh, A., and Mirasso, C · 2021
Cited alongside, same era.
Scaling vision with sparse mixture of experts
Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Pinto, A. S., Keysers, D., and Houlsby, N · 2021
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with GPT-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S. M., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y · 2023
Later among the works it cites.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J., Sifre, L., and Jumper, J · 2023
Later among the works it cites.
Optimize weight rounding via signed gradient descent for the quantization of llms
Cheng, W., Zhang, W., Shen, H., Cai, Y., He, X., and Lv, K · 2023
Later among the works it cites.
GPTQ: Accurate quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
So, D. R., Manke, W., Liu, H., Dai, Z., Shazeer, N., and Le, Q. V · 2021
Cited alongside, same era.
Lightseq: A high performance inference library for transformers
Wang, X., Xiong, Y., Wei, Y., Wang, M., and Li, L · 2021
Cited alongside, same era.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2021
Cited alongside, same era.
Deepspeed- inference: Enabling efficient inference of transformer models at unprecedented scale
Aminabadi, R. Y., Rajbhandari, S., Awan, A. A., Li, C., Li, D., Zheng, E., Ruwase, O., Smith, S., Zhang, M., Rasley, J., and He, Y · 2022
Cited alongside, same era.
Efficient large scale language modeling with mixtures of experts
Artetxe, M., Bhosale, S., Goyal, N., Mihaylov, T., Ott, M., Shleifer, S., Lin, X. V., Du, J., Iyer, S., Pasunuru, R., Anantharaman, G., Li, X., Chen, S., Akin, H., Baines, M., Martin, L., Zhou, X., Koura, P. S., O’Horo, B., Wang, J., Zettlemoyer, L., Diab, M., Kozareva, Z., and Stoyanov, V · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., Zoph, B., Fedus, L., Bosma, M. P., Zhou, Z., Wang, T., Wang, Y. E., Webster, K., Pellat, M., Robinson, K., Meier-Hellstern, K. S., Duke, T., Dixon, L., Zhang, K., Le, Q. V., Wu, Y., Chen, Z., and Cui, C · 2022
Cited alongside, same era.
Gu, A. and Dao, T · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Later among the works it cites.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Later among the works it cites.
The lazy neuron phenomenon: On emergence of activation sparsity in transformers
Li, Z., You, C., Bhojanapalli, S., Li, D., Rawat, A. S., Reddi, S. J., Ye, K., Chern, F., Yu, F., Guo, R., and Kumar, S · 2023
Later among the works it cites.
AWQ: activation-aware weight quantization for LLM compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Later among the works it cites.
Deja vu: Contextual sparsity for efficient llms at inference time
Liu, Z., Wang, J., Dao, T., Zhou, T., Yuan, B., Song, Z., Shrivastava, A., Zhang, C., Tian, Y., Ré, C., and Chen, B · 2023
Later among the works it cites.
Relu strikes back: Exploiting activation sparsity in large language models
Mirzadeh, I., Alizadeh, K., Mehta, S., Mundo, C. C. D., Tuzel, O., Samei, G., Rastegari, M., and Farajtabar, M · 2023
Later among the works it cites.
Llm foundry
Mosaicml · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., and Launay, J · 2023
Later among the works it cites.
RWKV: reinventing rnns for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Derczynski, L., Du, X., Grella, M., Gv, K., He, X., Hou, H., Kazienko, P., Kocon, J., Kong, J., Koptyra, B., Lau, H., Lin, J., Mantri, K. S. I., Mom, F., Saito, A., Song, G., Tang, X., Wind, J. S., Wozniak, S., Zhang, Z., Zhou, Q., Zhu, J., and Zhu, R · 2023
Later among the works it cites.
Flexgen: High-throughput generative inference of large language models with a single GPU
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Later among the works it cites.
SlimPajama: A 627B token cleaned and deduplicated version of RedPajama, 2023
Soboleva, D., Al-Khateeb, F., Myers, R., Steeves, J. R., Hestness, J., and Dey, N · 2023
Later among the works it cites.
Powerinfer: Fast large language model serving with a consumer-grade gpu
Song, Y., Mi, Z., Xie, H., and Chen, H · 2023
Later among the works it cites.
Accelerating LLM inference with staged speculative decoding
Spector, B. and Ré, C · 2023
Later among the works it cites.
Algorithm and hardness for dynamic attention maintenance in large language models
van den Brand, J., Song, Z., and Zhou, T · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Later among the works it cites.
Emergent modularity in pre-trained transformers
Zhang, Z., Zeng, Z., Lin, Y., Xiao, C., Wang, X., Han, X., Liu, Z., Xie, R., Sun, M., and Zhou, J · 2023
Later among the works it cites.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., Lee, J. D., Chen, D., and Dao, T · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Closest in time.