Fetching the paper…
Reading the bibliography…
Autoregressive Large Language Models (LLMs) have achieved impressive performance in language tasks but face two significant bottlenecks: (1) quadratic complexity in the attention module as the number of tokens increases, and (2) limited efficiency due to the sequential processing nature of autoregressive LLMs during generation.
Building a Large Annotated Corpus of English: The Penn Treebank
Marcus, M., Santorini, B., and Marcinkiewicz, M. A · 1993
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Automatically Constructing a Corpus of Sentential Paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
The PASCAL Recognising Textual Entailment Challenge
Dagan, I., Glickman, O., and Magnini, B · 2006
Earlier work this paper cites.
The Human Knowledge Compression Contest
Hutter, M · 2012
Earlier work this paper cites.
The winograd Schema Challenge
Levesque, H., Davis, E., and Morgenstern, L · 2012
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Character-level Convolutional Networks for Text Classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Quora Question Pairs, 2017
DataCanary, hilfialkaff, Jiang, L., Risdal, M., Dandekar, N., and tomtung · 2017
Earlier work this paper cites.
Language Modeling with Gated Convolutional Networks
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D · 2017
Earlier work this paper cites.
Pointer Sentinel Mixture Models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Earlier work this paper cites.
Get To The Point: Summarization with Pointer-Generator Networks
See, A., Liu, P. J., and Manning, C. D · 2017
Earlier work this paper cites.
Parallel Multi Channel Convolution using General Matrix Multiplication
Vasudevan, A., Anderson, A., and Gregg, D · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Think You Have Solved Question Answering? Try Arc, the AI2 Reasoning Challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
GLUE: A Multi-task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A., Nangia, N., and Bowman, S · 2018
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Compressive Transformers for Long-range Sequence Modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., and Lillicrap, T. P · 2019
Earlier work this paper cites.
Energy and Policy Considerations for Deep Learning in NLP
Strubell, E., Ganesh, A., and McCallum, A · 2019
Earlier work this paper cites.
SuperGLUE: A Stickier Benchmark for General-purpose Language Understanding Systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2019
Earlier work this paper cites.
PIQA: Reasoning about Physical Commonsense in Natural Language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Cited alongside, same era.
Language Models are Few-shot Learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Wiki-40B: Multilingual Language Model Dataset
Guo, M., Dai, Z., Vrandečić, D., and Al-Rfou, R · 2020
Cited alongside, same era.
Measuring Massive Multitask Language Understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Cited alongside, same era.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Cited alongside, same era.
Confident Adaptive Language Modeling
Schuster, T., Fisch, A., Gupta, J., Dehghani, M., Bahri, D., Tran, V., Tay, Y., and Metzler, D · 2022
Later among the works it cites.
Challenging Big-Bench Tasks and Whether Chain-of-Thought can Solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al · 2022
Later among the works it cites.
Maxvit: Multi-axis Vision Transformer
Tu, Z., Talebi, H., Zhang, H., Yang, F., Milanfar, P., Bovik, A., and Li, Y · 2022
Later among the works it cites.
Gemini: A Family of Highly Capable Multimodal Models
Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Bae, S., Ko, J., Song, H., and Yun, S.-Y · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Linformer: Self-Attention with Linear Complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Cited alongside, same era.
Rethinking Attention with Performers
Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., Belanger, D. B., Colwell, L. J., and Weller, A · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Cited alongside, same era.
SOFT: Softmax-free Transformer with Linear Complexity
Lu, J., Yao, J., Zhang, J., Zhu, X., Xu, H., Gao, W., Xu, C., Xiang, T., and Zhang, L · 2021
Cited alongside, same era.
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N. A., and Kong, L · 2021
Cited alongside, same era.
Later among the works it cites.
RedPajama: An Open Dataset for Training Large Language Models, October 2023
Computer, T · 2023
Later among the works it cites.
FLAT: An Optimized Dataflow for Mitigating Attention Bottlenecks
Kao, S.-C., Subramanian, S., Agrawal, G., Yazdanbakhsh, A., and Krishna, T · 2023
Later among the works it cites.
Big Little Transformer Decoder
Kim, S., Mangalam, K., Malik, J., Mahoney, M. W., Gholami, A., and Keutzer, K · 2023
Later among the works it cites.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Later among the works it cites.
SEA: Sparse Linear Attention with Estimated Attention Mask
Lee, H., Kim, J., Willette, J., and Hwang, S. J · 2023
Later among the works it cites.
Fast Inference from Transformers via Speculative Decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Later among the works it cites.
LLM-Pruner: On the Structural Pruning of Large Language Models
Ma, X., Fang, G., and Wang, X · 2023
Later among the works it cites.
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Wong, R. Y. Y., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z · 2023
Later among the works it cites.
Scaling Transnormer to 175 Billion Parameters
Qin, Z., Li, D., Sun, W., Sun, W., Shen, X., Han, X., Wei, Y., Lv, B., Yuan, F., Luo, X., et al · 2023
Later among the works it cites.
Stanford Alpaca: An Instruction-Following LLaMA Model, 2023
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Google’s AI chatbot “Bard”: A Side-by-Side Comparison with ChatGPT and its Utilization in Ophthalmology
Waisberg, E., Ong, J., Masalkhi, M., Zaman, N., Sarker, P., Lee, A. G., and Tavakkoli, A · 2023
Later among the works it cites.
SmoothQuant: Accurate and Efficient Post-training Quantization for Large Language Models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Later among the works it cites.
Gated Linear Attention Transformers with Hardware-efficient Training
Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y · 2023
Later among the works it cites.
ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design
You, H., Sun, Z., Shi, H., Yu, Z., Zhao, Y., Zhang, Y., Li, C., Li, B., and Lin, Y · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., Tumanov, A., and Ramjee, R · 2024
Closest in time.
Effective Interplay between Sparsity and Quantization: From Theory to Practice
Harma, S. B., Chakraborty, A., Kostenok, E., Mishin, D., Ha, D., Falsafi, B., Jaggi, M., Liu, M., Oh, Y., Subramanian, S., and Yazdanbakhsh, A · 2024
Closest in time.
ShiftAddViT: Mixture of Multiplication Primitives Towards Efficient Vision Transformer
You, H., Shi, H., Guo, Y., and Lin, Y · 2024
Closest in time.