Fetching the paper…
Reading the bibliography…
Linear position interpolation helps pre-trained models using rotary position embeddings (RoPE) to extrapolate to longer sequence lengths.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2010
Earlier work this paper cites.
Efficient Attentions for Long Document Summarization
Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, and Lu Wang · 2021
Earlier work this paper cites.
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, 2021
Ofir Press, Noah A. Smith, and Mike Lewis · 2021
Earlier work this paper cites.
QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization
Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, and Dragomir Radev · 2021
Earlier work this paper cites.
Proof-pile, 2022
Zhangir Azerbayev, Edward Ayers, and Bartosz Piotrowski · 2022
Cited alongside, same era.
An Empirical Analysis of Compute-optimal Large Language Model Training
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
RoFormer: Enhanced Transformer with Rotary Position Embedding, 2022
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2022
Cited alongside, same era.
Extending Context Window of Large Language Models via Positional Interpolation, 2023
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian · 2023
Cited alongside, same era.
Announcing MPT-7B-8K: 8K Context Length for Document Understanding, 2023a
MosaicML NLP Team
Cited in the paper.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023b
MosaicML NLP Team
Cited in the paper.
BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model, 2023
Nolan Dey*, Daria Soboleva*, Faisal Al-Khateeb, Bowen Yang, Ribhu Pathria, Hemant Khachane, Shaheer Muhammad, Zhiming (Charles) Chen, Robert Myers, Jacob Robert Steeves, et al · 2023
Closest in time.
Things I’m learning while training superhot, 2023
kaiokendev · 2023
Closest in time.
How Long Can Open-Source LLMs Truly Promise on Context Length?, 2023
Dacheng Li*, Rulin Shao*, Anze Xie, Ying Sheng, Lianmin Zheng, Joseph E. Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang · 2023
Closest in time.
YaRN: Efficient Context Window Extension of Large Language Models, 2023
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…