Fetching the paper…
Reading the bibliography…
The Transformer architecture has significantly advanced natural language processing (NLP) and has been foundational in developing large language models (LLMs) such as LLaMA and OPT, which have come to dominate a broad range of NLP tasks.
L. A. Belady, R. A. Nelson, and G. S. Shedler, “An anomaly in space-time characteristics of certain programs running in a paging machine,” Commun. ACM , vol. 12, pp. 349–353, 1969
1969
Earlier work this paper cites.
M. P. Marcus, B. Santorini, and M. A. Marcinkiewicz, “Building a large annotated corpus of english: The penn treebank,” Comput. Linguistics , vol. 19, pp. 313–330, 1993
1993
Earlier work this paper cites.
Y. Guo, A. Yao, and Y. Chen, “Dynamic network surgery for efficient dnns,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
W. Wen, C. Wu, Y. Wang, Y. Chen, and H. Li, “Learning structured sparsity in deep neural networks,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Yeo, G. Lee, G. Wang, S. Choi, H. Cho, R. K. Amplayo, and S. won Hwang, “Visual choice of plausible alternatives: An evaluation of image-based commonsense causal reasoning,” in Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) , 2018
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , vol. 1, 2019, p. 2
2019
Earlier work this paper cites.
N. Kitaev, L. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” International Conference on Learning Representations (ICLR) , 2019
2019
Earlier work this paper cites.
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” in North American Chapter of the Association for Computational Linguistics , 2019, pp. 6151–6162
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” 2019. [Online]. Available: https://insightcivic.s3.us-east-1.amazonaws.com/language-models.pdf
2019
Earlier work this paper cites.
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi, “Winogrande: An adversarial winograd schema challenge at scale,” Commun. ACM , vol. 64, pp. 99–106, 2019
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, and et al, “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Cited alongside, same era.
B. Chmiel, R. Banner, G. Shomron, Y. Nahshan, A. Bronstein, U. Weiser et al. , “Robust quantization: One model to rule them all,” Advances in neural information processing systems , vol. 33, pp. 5308–5317, 2020
2020
Cited alongside, same era.
2022
Later among the works it cites.
2023
Later among the works it cites.
S. Dai, H. Genc, R. Venkatesan, and B. Khailany, “Efficient transformer inference with statically structured sparse attention,” in 2023 60th ACM/IEEE Design Automation Conference (DAC) . IEEE, 2023, pp. 1–6
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) , pp. 97–110, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, and J. Brew, “Transformers: State-of-the-art natural language processing,” in Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations , 2020, pp. 38–45
2020
Cited alongside, same era.
B. Chen, Z. Liu, B. Peng, Z. Xu, J. L. Li, T. Dao, Z. Song, A. Shrivastava, and C. Ré, “Mongoose: A learnable lsh framework for efficient neural network training,” International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
L. Gao, J. Tow, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, K. McDonell, N. Muennighoff, J. Phang, L. Reynolds, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou, “A framework for few-shot language model evaluation,” 2021. [Online]. Available: https://doi.org/10.5281/zenodo.5371628
2021
Cited alongside, same era.
H. Shi, J. Gao, X. Ren, H. Xu, X. Liang, Z. Li, and J. T.-Y. Kwok, “Sparsebert: Rethinking the importance analysis in self-attention,” in International Conference on Machine Learning . PMLR, 2021, pp. 9547–9557
2021
Cited alongside, same era.
R. Y. Aminabadi, S. Rajbhandari, M. Zhang, A. A. Awan, C. Li, D. Li, E. Zheng, J. Rasley, S. Smith, O. Ruwase, and Y. He, “Deepspeed- inference: Enabling efficient inference of transformer models at unprecedented scale,” SC22: International Conference for High Performance Computing, Networking, Storage and Analysis , pp. 1–15, 2022
2022
Cited alongside, same era.
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” Advances in Neural Information Processing Systems , vol. 35, pp. 16 344–16 359, 2022
2022
Cited alongside, same era.
T. Dettmers and L. Zettlemoyer, “The case for 4-bit precision: k-bit inference scaling laws,” in International Conference on Machine Learning . PMLR, 2023, pp. 7750–7774
2023
Later among the works it cites.
J. Ding, S. Ma, L. Dong, X. Zhang, S. Huang, W. Wang, N. Zheng, and F. Wei, “Longnet: Scaling transformers to 1,000,000,000 tokens,” ArXiv , vol. abs/22307.02486, 2023
2023
Later among the works it cites.
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh, “Gptq: Accurate post-training quantization for generative pre-trained transformers,” International Conference on Learning Representations (ICLR) , 2023
2023
Later among the works it cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , pp. 611–626, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, J. Heek, K. Xiao, S. Agrawal, and J. Dean, “Efficiently scaling transformer inference,” Proceedings of Machine Learning and Systems , vol. 5, 2023
2023
Later among the works it cites.
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, D. Y. Fu, Z. Xie, B. Chen, C. W. Barrett, J. Gonzalez, P. Liang, C. Ré, I. C. Stoica, and C. Zhang, “High-throughput generative inference of large language models with a single gpu,” in International Conference on Machine Learning . PMLR, 2023, pp. 31 094––31 116
2023
Later among the works it cites.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto, “Stanford alpaca: An instruction-following llama model,” https://github.com/tatsu-lab/stanford_alpaca , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.