Fetching the paper…
Reading the bibliography…
Recent studies have shown that Large Language Models (LLMs) have the potential to process extremely long text.
The narrativeqa reading comprehension challenge, 2017
Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette · 2017
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling, 2019
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, and Timothy P. Lillicrap · 2019
Earlier work this paper cites.
Qmsum: A new benchmark for query-based multi-domain meeting summarization, 2021
Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, and Dragomir Radev · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation, 2022
Ofir Press, Noah A. Smith, and Mike Lewis · 2022
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding, 2023
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li · 2023
Cited alongside, same era.
Longnet: Scaling transformers to 1,000,000,000 tokens, 2023
Jiayu Ding, Shuming Ma, Li Dong, Xingxing Zhang, Shaohan Huang, Wenhui Wang, Nanning Zheng, and Furu Wei · 2023
Cited alongside, same era.
How long can open-source llms truly promise on context length?, June 2023
Dacheng Li, Rulin Shao, Anze Xie, Ying Sheng, Lianmin Zheng, Joseph E. Gonzalez, Ion Stoica, Xuezhe Ma, , and Hao Zhang · 2023
Cited alongside, same era.
Clex: Continuous length extrapolation for large language models, 2023a
Guanzheng Chen, Xin Li, Zaiqiao Meng, Shangsong Liang, and Lidong Bing
Cited in the paper.
Extending context window of large language models via positional interpolation, 2023b
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian
Cited in the paper.
Longlora: Efficient fine-tuning of long-context large language models, 2023c
Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia
Cited in the paper.
Scaling laws of rope-based extrapolation, 2023
Xiaoran Liu, Hang Yan, Shuo Zhang, Chenxin An, Xipeng Qiu, and Dahua Lin · 2023
Later among the works it cites.
Yarn: Efficient context window extension of large language models, 2023
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole · 2023
Later among the works it cites.
Effective long-context scaling of foundation models, 2023
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, and Hao Ma · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…