Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have exhibited exceptional performance across a spectrum of natural language processing tasks.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Pointer Sentinel Mixture Models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Crowdsourcing Multiple Choice Science Questions
Welbl, J., Liu, N., and Gardner, M · 2017
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Beltagy, I., Peters, M., and Cohan, A · 2020
Earlier work this paper cites.
PIQA: Reasoning about Physical Commonsense in Natural Language
Bisk, Y., Zellers, R., Bras, R., Gao, J., and Choi, Y · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2020
Earlier work this paper cites.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Earlier work this paper cites.
Compressive Transformers for Long-Range Sequence Modelling
Rae, J., Potapenko, A., Jayakumar, S., Hillier, C., and Lillicrap, T · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P · 2020
Earlier work this paper cites.
O(n) connections are expressive enough: Universal approximability of sparse transformers
Yun, C., Chang, Y.-W., Bhojanapalli, S., Rawat, A. S., Reddi, S., and Kumar, S · 2020
Cited alongside, same era.
Big Bird: Transformers for Longer Sequences
Zaheer, M., Guruganesh, G., Dubey, K., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A · 2020
Cited alongside, same era.
PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization
Zhang, J., Zhao, Y., Saleh, M., and Liu, P · 2020
Cited alongside, same era.
Training Verifiers to Solve Math Word Problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Cited alongside, same era.
Measuring Massive Multitask Language Understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
Chen, Y., Qian, S., Tang, H., Lai, X., Liu, Z., Han, S., and Jia, J · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J · 2024
Closest in time.
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
He, Z., Feng, G., Luo, S., Yang, K., Wang, L., Xu, J., Zhang, Z., Yang, H., and He, D · 2024
Closest in time.
SnapKV: LLM Knows What You Are Looking for Before Generation
Li, Y., Huang, Y., Yang, B., Venkitesh, B., Locatelli, A., Ye, H., Cai, T., Lewis, P., and Chen, D · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning
Liu, J., Cui, L., Liu, H., Huang, D., Wang, Y., and Zhang, Y · 2021
Cited alongside, same era.
Linear transformers are secretly fast weight programmers
Schlag, I., Irie, K., and Schmidhuber, J · 2021
Cited alongside, same era.
SparseBERT: Rethinking the Importance Analysis in Self-attention
Shi, H., Gao, J., Ren, X., Xu, H., Liang, X., Li, Z., and Kwok, J · 2021
Cited alongside, same era.
Transformers for Modeling Physical Systems
Geneva, N. and Zabaras, N · 2022
Cited alongside, same era.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q., Zhou, D., et al · 2022
Cited alongside, same era.
The Falcon Series of Open Language Models
Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Debbah, M., Goffinet, É., Hesslow, D., Launay, J., Malartic, Q., Mazzotta, D., Noune, B., Pannier, B., and Penedo, G · 2023
Cited alongside, same era.
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Biderman, S., Schoelkopf, H., Anthony, Q., Bradley, H., O’Brien, K., Hallahan, E., Khan, M., Purohit, S., Prashanth, U., Raff, E., Skowron, A., Sutawika, L., and Wal, O · 2023
Cited alongside, same era.
Flexattention: The flexibility of pytorch with the performance of flashattention, 2024
PyTorch · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Closest in time.
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
Yang, D., Han, X., Gao, Y., Hu, Y., Zhang, S., and Zhao, H · 2024
Closest in time.
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Zhang, Y., Gao, B., Liu, T., Lu, K., Xiong, W., Dong, Y., Chang, B., Hu, J., Xiao, W., et al · 2024
Closest in time.
Zhu, Q., Duan, J., Chen, C., Liu, S., Li, X., Feng, G., Lv, X., Cao, H., Xiao, C., Zhang, X., et al · 2024
Closest in time.
QuickLLaMA: Query-Aware Inference Acceleration for Large Language Models
Li, J., Shi, H., Wu, S., Zheng, C., Li, Z., Jiang, X., Xu, H., and Jia, J · 2025
Closest in time.
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Yuan, J., Gao, H., Dai, D., Luo, J., Zhao, L., Zhang, Z., Xie, Z., Wei, Y. X., Wang, L., Xiao, Z., Wang, Y., Ruan, C., Zhang, M., Liang, W., and Zeng, W · 2025
Closest in time.
Self-adjust softmax
Zheng, C., Gao, Y., Chen, G., Shi, H., Xiong, J., Ren, X., Huang, C., Jiang, X., Li, Z., and Li, Y · 2025
Closest in time.