Fetching the paper…
Reading the bibliography…
The inference demand for LLMs has skyrocketed in recent months, and serving models with low latencies remains challenging due to the quadratic input length complexity of the attention layers.
Adaptive computation time for recurrent neural networks, 2017
Graves, A · 2017
Earlier work this paper cites.
Branchynet: Fast inference via early exiting from deep neural networks, 2017
Teerapittayanon, S., McDanel, B., and Kung, H. T · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout, 2019
Fan, A., Grave, E., and Joulin, A · 2019
Earlier work this paper cites.
Efficient training of bert by progressively stacking
Gong, L., He, D., Li, Z., Qin, T., Wang, L., and Liu, T · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?, 2019
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Depth-adaptive transformer
Elbayad, M., Gu, J., Grave, E., and Auli, M · 2020
Earlier work this paper cites.
Learning layer-skippable inference network
Jiang, Y.-G., Cheng, C., Lin, H., and Fu, Y · 2020
Earlier work this paper cites.
Accelerating training of transformer-based language models with progressive layer dropping
Zhang, M. and He, Y · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Carbon emissions and large neural network training
Patterson, D. A., Gonzalez, J., Le, Q. V., Liang, C., Munguia, L., Rothchild, D., So, D. R., Texier, M., and Dean, J · 2021
Earlier work this paper cites.
Hungry hungry hippos: Towards language modeling with state space models
Fu, D. Y., Dao, T., Saab, K. K., Thomas, A. W., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods, 2022
Lin, S., Hilton, J., and Evans, O · 2022
Earlier work this paper cites.
On the effect of dropping layers of pre-trained transformer models
Sajjad, H., Dalvi, F., Durrani, N., and Nakov, P · 2022
Earlier work this paper cites.
Confident adaptive language modeling
Schuster, T., Fisch, A., Gupta, J., Dehghani, M., Bahri, D., Tran, V., Tay, Y., and Metzler, D · 2022
Cited alongside, same era.
Centered self-attention layers, 2023
Ali, A., Galanti, T., and Wolf, L · 2023
Cited alongside, same era.
Open llm leaderboard
Beeching, E., Fourrier, C., Habib, N., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T · 2023
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens
Belrose, N., Furman, Z., Smith, L., Halawi, D., Ostrovsky, I., McKinney, L., Biderman, S., and Steinhardt, J · 2023
Cited alongside, same era.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J., Sifre, L., and Jumper, J · 2023
Cited alongside, same era.
Stabilizing transformer training by preventing attention entropy collapse, 2023
Zhai, S., Likhomanenko, T., Littwin, E., Busbridge, D., Ramapuram, J., Zhang, Y., Gu, J., and Susskind, J · 2023
Later among the works it cites.
Deepseek llm: Scaling open-source language models with longtermism
Bi, X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., Ding, H., Dong, K., Du, Q., Fu, Z., et al · 2024
Closest in time.
Ee-llm: Large-scale training and inference of early-exit large language models with 3d parallelism, 2024
Chen, Y., Pan, X., Li, Y., Ding, B., and Zhou, J · 2024
Closest in time.
Graph convolutions enrich the self-attention in transformers!, 2024
Choi, J., Wi, H., Kim, J., Shin, Y., Lee, K., Trask, N., and Park, N · 2024
Closest in time.
Setting the record straight on transformer oversmoothing, 2024
Dovonon, G. J.-S., Bronstein, M. M., and Kusner, M. J · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Din, A. Y., Karidi, T., Choshen, L., and Geva, M · 2023
Cited alongside, same era.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth, 2023
Dong, Y., Cordonnier, J.-B., and Loukas, A · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces, 2023
Gu, A. and Dao, T · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
No train no gain: Revisiting efficient training algorithms for transformer-based language models
Kaddour, J., Key, O., Nawrot, P., Minervini, P., and Kusner, M. J · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Deja vu: Contextual sparsity for efficient LLMs at inference time
Liu, Z., Wang, J., Dao, T., Zhou, T., Yuan, B., Song, Z., Shrivastava, A., Zhang, C., Tian, Y., Re, C., and Chen, B · 2023
Cited alongside, same era.
Elhoushi, M., Shrivastava, A., Liskovich, D., Hosmer, B., Wasti, B., Lai, L., Mahmoud, A., Acun, B., Agarwal, S., Roman, A., et al · 2024
Closest in time.
Not all layers of llms are necessary during inference, 2024
Fan, S., Jiang, X., Li, X., Meng, X., Han, P., Shang, S., Sun, A., Wang, Y., and Wang, Z · 2024
Closest in time.
The unreasonable ineffectiveness of the deeper layers, 2024
Gromov, A., Tirumala, K., Shapourian, H., Glorioso, P., and Roberts, D. A · 2024
Closest in time.
Shortgpt: Layers in large language models are more redundant than you expect, 2024
Men, X., Xu, M., Zhang, Q., Wang, B., Lin, H., Lu, Y., Han, X., and Chen, W · 2024
Closest in time.
Dynamic memory compression: Retrofitting llms for accelerated inference, 2024
Nawrot, P., Łańcucki, A., Chochowski, M., Tarjan, D., and Ponti, E. M · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Closest in time.
Layer-condensed kv cache for efficient inference of large language models, 2024
Wu, H. and Tu, K · 2024
Closest in time.
Xia, H., Yang, Z., Dong, Q., Wang, P., Li, Y., Ge, T., Liu, T., Li, W., and Sui, Z · 2024
Closest in time.