Fetching the paper…
Reading the bibliography…
The rapid expansion of context window sizes in Large Language Models~(LLMs) has enabled them to tackle increasingly complex tasks involving lengthy documents.
Language Models are Few-Shot Learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 1901
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Shazeer, N. 2019 · 1911
Earlier work this paper cites.
The Jensen-Shannon divergence
Menéndez, M.; Pardo, J.; Pardo, L.; and Pardo, M. 1997 · 1997
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Beltagy, I.; Peters, M. E.; and Cohan, A. 2020 · 2004
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019 · 2019
Earlier work this paper cites.
Sharing Attention Weights for Fast Transformer
Xiao, T.; Li, Y.; Zhu, J.; Yu, Z.; and Liu, T. 2019 · 2019
Earlier work this paper cites.
ZeRO: memory optimizations toward training trillion parameter models
Rajbhandari, S.; Rasley, J.; Ruwase, O.; and He, Y. 2020 · 2020
Earlier work this paper cites.
DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters
Rasley, J.; Rajbhandari, S.; Ruwase, O.; and He, Y. 2020 · 2020
Earlier work this paper cites.
Transformers: State-of-the-Art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Le Scao, T.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. 2020 · 2020
Earlier work this paper cites.
Leveraging redundancy in attention with reuse transformers
Bhojanapalli, S.; Chakrabarti, A.; Veit, A.; Lukasik, M.; Jain, H.; Liu, F.; Chang, Y.-W.; and Kumar, S. 2021 · 2021
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Earlier work this paper cites.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Ainslie, J.; Lee-Thorp, J.; de Jong, M.; Zemlyanskiy, Y.; Lebron, F.; and Sanghai, S. 2023 · 2023
Earlier work this paper cites.
Jacobs, S. A.; Tanaka, M.; Zhang, C.; Zhang, M.; Song, S. L.; Rajbhandari, S.; and He, Y. 2023 · 2023
Cited alongside, same era.
Compressing Context to Enhance Inference Efficiency of Large Language Models
Li, Y.; Dong, B.; Guerin, F.; and Lin, C. 2023 · 2023
Cited alongside, same era.
Efficiently scaling transformer inference
Pope, R.; Douglas, S.; Chowdhery, A.; Devlin, J.; Bradbury, J.; Heek, J.; Xiao, K.; Agrawal, S.; and Dean, J. 2023 · 2023
Cited alongside, same era.
FlexGen: high-throughput generative inference of large language models with a single GPU
Sheng, Y.; Zheng, L.; Yuan, B.; Li, Z.; Ryabinin, M.; Chen, B.; Liang, P.; Ré, C.; Stoica, I.; and Zhang, C. 2023 · 2023
Cited alongside, same era.
RoFormer: Enhanced Transformer with Rotary Position Embedding
Su, J.; Lu, Y.; Pan, S.; Murtadha, A.; Wen, B.; and Liu, Y. 2023 · 2023
How to Train Long-Context Language Models (Effectively)
Gao, T.; Wettig, A.; Yen, H.; and Chen, D. 2024 · 2024
Closest in time.
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
GLM, T.; Zeng, A.; Xu, B.; Wang, B.; Zhang, C.; Yin, D.; Rojas, D.; Feng, G.; Zhao, H.; Lai, H.; et al. 2024 · 2024
Closest in time.
LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models
Han, C.; Wang, Q.; Peng, H.; Xiong, W.; Chen, Y.; Ji, H.; and Wang, S. 2024 · 2024
Closest in time.
Snapkv: Llm knows what you are looking for before generation
Li, Y.; Huang, Y.; Yang, B.; Venkitesh, B.; Locatelli, A.; Ye, H.; Cai, T.; Lewis, P.; and Chen, D. 2024 · 2024
Closest in time.
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Cited alongside, same era.
H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
Zhang, Z.; Sheng, Y.; Zhou, T.; Chen, T.; Zheng, L.; Cai, R.; Song, Z.; Tian, Y.; Re, C.; Barrett, C.; Wang, Z.; and Chen, B. 2023 · 2023
Cited alongside, same era.
L-Eval: Instituting Standardized Evaluation for Long Context Language Models
An, C.; Gong, S.; Zhong, M.; Zhao, X.; Li, M.; Zhang, J.; Kong, L.; and Qiu, X. 2024 · 2024
Cited alongside, same era.
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Bai, Y.; Lv, X.; Zhang, J.; Lyu, H.; Tang, J.; Huang, Z.; Du, Z.; Liu, X.; Zeng, A.; Hou, L.; Dong, Y.; Tang, J.; and Li, J. 2024 · 2024
Cited alongside, same era.
Codeplan: Repository-level coding using llms and planning
Bairi, R.; Sonwane, A.; Kanade, A.; Iyer, A.; Parthasarathy, S.; Rajamani, S.; Ashok, B.; and Shet, S. 2024 · 2024
Cited alongside, same era.
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Brandon, W.; Mishra, M.; Nrusimha, A.; Panda, R.; and Kelly, J. R. 2024 · 2024
Cited alongside, same era.
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024 · 2024
Cited alongside, same era.
Lin, J.; Tang, J.; Tang, H.; Yang, S.; Chen, W.-M.; Wang, W.-C.; Xiao, G.; Dang, X.; Gan, C.; and Han, S. 2024 · 2024
Closest in time.
Lifelong and Continual Learning Dialogue Systems
Mazumder, S.; and Liu, B. 2024 · 2024
Closest in time.
LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression
Pan, Z.; Wu, Q.; Jiang, H.; Xia, M.; Luo, X.; Zhang, J.; Lin, Q.; Rühle, V.; Yang, Y.; Lin, C.-Y.; Zhao, H. V.; Qiu, L.; and Zhang, D. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M.; Savinov, N.; Teplyashin, D.; Lepikhin, D.; Lillicrap, T.; Alayrac, J.-b.; Soricut, R.; Lazaridou, A.; Firat, O.; Schrittwieser, J.; et al. 2024 · 2024
Closest in time.
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Soldaini, L.; Kinney, R.; Bhagia, A.; Schwenk, D.; Atkinson, D.; Authur, R.; Bogin, B.; Chandu, K.; Dumas, J.; Elazar, Y.; Hofmann, V.; Jha, A.; Kumar, S.; Lucy, L.; Lyu, X.; Lambert, N.; Magnusson, I.; Morrison, J.; Muennighoff, N.; Naik, A.; Nam, C.; Peters, M.; Ravichander, A.; Richardson, K.; Shen, Z.; Strubell, E.; Subramani, N.; Tafjord, O.; Walsh, E.; Zettlemoyer, L.; Smith, N.; Hajishirzi, H.; Beltagy, I.; Groeneveld, D.; Dodge, J.; and Lo, K. 2024 · 2024
Closest in time.
QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Tang, J.; Zhao, Y.; Zhu, K.; Xiao, G.; Kasikci, B.; and Han, S. 2024 · 2024
Closest in time.
Layer-Condensed KV Cache for Efficient Inference of Large Language Models
Wu, H.; and Tu, K. 2024 · 2024
Closest in time.
Efficient Streaming Language Models with Attention Sinks
Xiao, G.; Tian, Y.; Chen, B.; Han, S.; and Lewis, M. 2024 · 2024
Closest in time.
Effective Long-Context Scaling of Foundation Models
Xiong, W.; Liu, J.; Molybog, I.; Zhang, H.; Bhargava, P.; Hou, R.; Martin, L.; Rungta, R.; Sankararaman, K. A.; Oguz, B.; Khabsa, M.; Fang, H.; Mehdad, Y.; Narang, S.; Malik, K.; Fan, A.; Bhosale, S.; Edunov, S.; Lewis, M.; Wang, S.; and Ma, H. 2024 · 2024
Closest in time.