Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) with hundreds of billions of parameters have transformed the field of machine learning.
Who belongs in the family?
Thorndike, R. L · 1953
Earlier work this paper cites.
8-bit approximations for parallelism in deep learning
Dettmers, T · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout
Fan, A., Grave, E., and Joulin, A · 2019
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Shazeer, N · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Voita, E., Talbot, D., Moiseev, F., Sennrich, R., and Titov, I · 2019
Earlier work this paper cites.
Structured pruning of large language models
Wang, Z., Wohlwend, J., and Lei, T · 2019
Earlier work this paper cites.
Drawing early-bird tickets: Towards more efficient training of deep networks
You, H., Li, C., Xu, P., Fu, Y., Wang, Y., Chen, X., Baraniuk, R. G., Wang, Z., and Lin, Y · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Aˆ 3: Accelerating attention mechanisms in neural networks with approximation
Ham, T. J., Jung, S. J., Kim, S., Oh, Y. H., Park, Y., Song, Y., Park, J.-H., Lee, S., Park, K., Lee, J. W., et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Cited alongside, same era.
When bert plays the lottery, all tickets are winning
Prasanna, S., Rogers, A., and Rumshisky, A · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Q-bert: Hessian based ultra low precision quantization of bert
Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
The optimal bert surgeon: Scalable and accurate second-order pruning for large language models
Kurtic, E., Campos, D., Nguyen, T., Frantar, E., Kurtz, M., Fineran, B., Goin, M., and Alistarh, D · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S · 2023
Later among the works it cites.
Efficient transformer inference with statically structured sparse attention
Dai, S., Genc, H., Venkatesan, R., and Khailany, B · 2023
Later among the works it cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Principal component networks: Parameter reduction early in training
Waleffe, R. and Rekatsinas, T · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Cited alongside, same era.
Elsa: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks
Ham, T. J., Lee, Y., Seo, S. H., Kim, S., Choi, H., Jung, S. J., and Lee, J. W · 2021
Cited alongside, same era.
The hardware lottery
Hooker, S · 2021
Cited alongside, same era.
I-bert: Integer-only bert quantization
Kim, S., Gholami, A., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Cited alongside, same era.
Block pruning for faster transformers
Lagunas, F., Charlaix, E., Sanh, V., and Rush, A. M · 2021
Cited alongside, same era.
Edgebert: Sentence-level energy optimizations for latency-aware multi-task nlp inference
Tambe, T., Hooper, C., Pentecost, L., Jia, T., Yang, E.-Y., Donato, M., Sanh, V., Whatmough, P., Rush, A. M., Brooks, D., et al · 2021
Cited alongside, same era.
Dao, T · 2023
Later among the works it cites.
The case for 4-bit precision: k-bit inference scaling laws
Dettmers, T. and Zettlemoyer, L · 2023
Later among the works it cites.
xformers - toolbox to accelerate research on transformers
facebookresearch · 2023
Later among the works it cites.
Hsieh, C.-Y., Li, C.-L., Yeh, C.-K., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C.-Y., and Pfister, T · 2023
Later among the works it cites.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Later among the works it cites.
Fact: Ffn-attention co-optimized transformer architecture with eager correlation prediction
Qin, Y., Wang, Y., Deng, D., Zhao, Z., Yang, X., Liu, L., Wei, S., Hu, Y., and Yin, S · 2023
Later among the works it cites.
On the effect of dropping layers of pre-trained transformer models
Sajjad, H., Dalvi, F., Durrani, N., and Nakov, P · 2023
Later among the works it cites.
Cuttlefish: Low-rank model training without all the tuning
Wang, H., Agarwal, S., Tanaka, Y., Xing, E., Papailiopoulos, D., et al · 2023
Later among the works it cites.
Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation
Xia, H., Ge, T., Wang, P., Chen, S.-Q., Wei, F., and Sui, Z · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Later among the works it cites.
Predictive pipelined decoding: A compute-latency trade-off for exact llm decoding
Yang, S., Lee, G., Cho, J., Papailiopoulos, D., and Lee, K · 2023
Later among the works it cites.
H _ 2 \_2 o: Heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., et al · 2023
Later among the works it cites.
Brainformers: Trading simplicity for efficiency
Zhou, Y., Du, N., Huang, Y., Peng, D., Lan, C., Huang, D., Shakeri, S., So, D., Dai, A. M., Lu, Y., et al · 2023
Later among the works it cites.
https://github.com/FMInference/DejaVu , 2024
Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time · 2024
Closest in time.