Fetching the paper…
Reading the bibliography…
Accommodating long sequences efficiently in autoregressive Transformers, especially within an extended context window, poses significant challenges due to the quadratic computational complexity and substantial KV memory requirements inherent in self-attention mechanisms.
Possible generalization of Boltzmann-Gibbs statistics
Constantino Tsallis · 1988
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma · 2006
Earlier work this paper cites.
Cluster-former: Clustering-based sparse transformer for long-range dependency encoding
Shuohang Wang, Luowei Zhou, Zhe Gan, Yen-Chun Chen, Yuwei Fang, Siqi Sun, Yu Cheng, and Jingjing Liu · 2009
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville · 2013
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
André F. T. Martins and Ramón Fernández Astudillo · 2016
Earlier work this paper cites.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Parallelizing linear recurrent neural nets over sequence length
Eric Martin and Chris Cundy · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam M. Shazeer, Alexander Ku, and Dustin Tran · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc V. Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2019
Earlier work this paper cites.
Axial attention in multidimensional transformers
Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans · 2019
Earlier work this paper cites.
Sparse sequence-to-sequence models
Ben Peters, Vlad Niculae, and André F. T. Martins · 2019
Earlier work this paper cites.
Blockwise self-attention for long document understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Scott Yih, Sinong Wang, and Jie Tang · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, and Timothy P. Lillicrap · 2019
Earlier work this paper cites.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamás Sarlós, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy J. Colwell, and Adrian Weller · 2020
Earlier work this paper cites.
Funnel-transformer: Filtering out sequential redundancy for efficient language processing
Zihang Dai, Guokun Lai, Yiming Yang, and Quoc V. Le · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Franccois Fleuret · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Earlier work this paper cites.
Do transformers need deep long-range memory?
Jack Rae and Ali Razavi · 2020
Cited alongside, same era.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Taghi Saffar, Ashish Vaswani, and David Grangier · 2020
Cited alongside, same era.
Sparse sinkhorn attention
Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan · 2020
Cited alongside, same era.
Fast transformers with clustered attention
Apoorv Vyas, Angelos Katharopoulos, and Franccois Fleuret · 2020
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Cited alongside, same era.
Dynamically Scaled RoPE further increases performance of long context LLaMA with zero fine-tuning, 2023
emozilla · 2023
Later among the works it cites.
A framework for few-shot language model evaluation, 12 2023
Leo Gao, Jonathan Tow, Baber Abbasi, et al · 2023
Later among the works it cites.
Model tells you what to discard: Adaptive kv cache compression for llms
Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Later among the works it cites.
Longcoder: A long-range pre-trained language model for code completion
Daya Guo, Canwen Xu, Nan Duan, Jian Yin, and Julian McAuley · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beidi Chen, Tri Dao, Eric Winsor, Zhao Song, Atri Rudra, and Chris Ré · 2021
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher R’e · 2021
Cited alongside, same era.
Memory-efficient transformers via top-k attention
Ankit Gupta, Guy Dar, Shaya Goodman, David Ciprut, and Jonathan Berant · 2021
Cited alongside, same era.
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A. Smith, and Lingpeng Kong · 2021
Cited alongside, same era.
Combiner: Full attention transformer with sparse computation cost
Hongyu Ren, Hanjun Dai, Zihang Dai, Mengjiao Yang, Jure Leskovec, Dale Schuurmans, and Bo Dai · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Cited alongside, same era.
Do long-range language models actually use long-range context?
Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, and Mohit Iyyer · 2021
Cited alongside, same era.
Chi Han, Qifan Wang, Hao Peng, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang · 2023
Later among the works it cites.
Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L’elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Later among the works it cites.
Conditional adapters: Parameter-efficient transfer learning with fast inference
Tao Lei, Junwen Bai, Siddhartha Brahma, Joshua Ainslie, Kenton Lee, Yanqi Zhou, Nan Du, Vincent Zhao, Yuexin Wu, Bo Li, Yu Zhang, and Ming-Wei Chang · 2023
Later among the works it cites.
Zichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang, Victor Xie, Zhaozhuo Xu, Anastasios Kyrillidis, and Anshumali Shrivastava · 2023
Later among the works it cites.
Landmark Attention: Random-Access Infinite Context Length for Transformers
Amirkeivan Mohtashami and Martin Jaggi · 2023
Later among the works it cites.
Resurrecting recurrent neural networks for long sequences
Antonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De · 2023
Later among the works it cites.
Faster causal attention over large sequences through sparse flash attention
Matteo Pagliardini, Daniele Paliotta, Martin Jaggi, and Franccois Fleuret · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, Quentin G. Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, G Kranthikiran, Xuming He, Haowen Hou, Przemyslaw Kazienko, Jan Kocoń, Jiaming Kong, Bartlomiej Koptyra, Hayden Lau, Krishna Sri Ipsit Mantri, Ferdinand Mom, Atsushi Saito, Xiangru Tang, Bolun Wang, Johan Sokrates Wind, Stansilaw Wozniak, Ruichong Zhang, Zhenyuan Zhang, Qihang Zhao, Peng Zhou, Jian Zhu, and Rui Zhu · 2023
Later among the works it cites.
SlimPajama: A 627B token cleaned and deduplicated version of RedPajama, 6 2023
Daria Soboleva, Faisal Al-Khateeb, Robert Myers, Jacob R Steeves, Joel Hestness, and Nolan Dey · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei · 2023
Later among the works it cites.
OpenAI Team · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Gated linear attention transformers with hardware-efficient training
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim · 2023
Later among the works it cites.
Trams: Training-free memory selection for long-range language modeling
Haofei Yu, Cunxiang Wang, Yue Zhang, and Wei Bi · 2023
Later among the works it cites.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu (Allen) Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark W. Barrett, Zhangyang Wang, and Beidi Chen · 2023
Later among the works it cites.
Data engineering for scaling language models to 128k context
Yao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue, Hanna Hajishirzi, Yoon Kim, and Hao Peng · 2024
Closest in time.
Megalodon: Efficient llm pretraining and inference with unlimited context length, 2024
Xuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen, Lili Yu, Hao Zhang, Jonathan May, Luke Zettlemoyer, Omer Levy, and Chunting Zhou · 2024
Closest in time.
Leave no context behind: Efficient infinite context transformers with infini-attention
Tsendsuren Munkhdalai, Manaal Faruqui, and Siddharth Gopal · 2024
Closest in time.
Transnormerllm: A faster and better large language model with improved transnormer, 2024
Zhen Qin, Dong Li, Weigao Sun, Weixuan Sun, Xuyang Shen, Xiaodong Han, Yunshen Wei, Baohong Lv, Xiao Luo, Yu Qiao, and Yiran Zhong · 2024
Closest in time.
Fla: A triton-based library for hardware-efficient implementations of linear attention mechanism, January 2024
Songlin Yang and Yu Zhang · 2024
Closest in time.
Lory: Fully differentiable mixture-of-experts for autoregressive language model pre-training
Zexuan Zhong, Mengzhou Xia, Danqi Chen, and Mike Lewis · 2024
Closest in time.