Fetching the paper…
Reading the bibliography…
Today's large language models (LLMs) typically train on short text segments (e.g., <4K tokens) due to the quadratic complexity of their Transformer architectures.
Empirical processes: Theory and applications
David Pollard. 1990 · 1990
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Learning to remember rare events
Lukasz Kaiser, Ofir Nachum, Aurko Roy, and Samy Bengio. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Decision theoretic generalizations of the pac model for neural net and other learning applications
David Haussler. 2018 · 2018
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2019 · 2019
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020 · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020 · 2020
Earlier work this paper cites.
Rethinking positional encoding in language pre-training
Guolin Ke, Di He, and Tie-Yan Liu. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. 2020 · 2020
Earlier work this paper cites.
Skyformer: Remodel self-attention with gaussian kernel and nyström method
Yifan Chen, Qi Zeng, Heng Ji, and Yun Yang. 2021 · 2021
Earlier work this paper cites.
A dataset of information-seeking questions and answers anchored in research papers
Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A Smith, and Matt Gardner. 2021 · 2021
Earlier work this paper cites.
Shape: Shifted absolute position embedding for transformers
Shun Kiyono, Sosuke Kobayashi, Jun Suzuki, and Kentaro Inui. 2021 · 2021
Earlier work this paper cites.
Cape: Encoding relative positions with continuous augmented positional embeddings
Tatiana Likhomanenko, Qiantong Xu, Gabriel Synnaeve, Ronan Collobert, and Alex Rogozhnikov. 2021 · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah Smith, and Mike Lewis. 2021 · 2021
Earlier work this paper cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2021 · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Cited alongside, same era.
Memorizing transformers
Yuhuai Wu, Markus Norman Rabe, DeLesley Hutchins, and Christian Szegedy. 2021 · 2021
Cited alongside, same era.
Adaptive semiparametric language models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong. 2021 · 2021
Cited alongside, same era.
Exploring length generalization in large language models
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur. 2022 · 2022
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022 · 2022
Cited alongside, same era.
Induced natural language rationales and interleaved markup tokens enable extrapolation in large language models
Zebra: Extending context window with layerwise grouped local-global attention
Kaiqiang Song, Xiaoyang Wang, Sangwoo Cho, Xiaoman Pan, and Dong Yu. 2023 · 2023
Closest in time.
A frustratingly easy improvement for position embeddings via random padding
Mingxu Tao, Yansong Feng, and Dongyan Zhao. 2023 · 2023
Closest in time.
Introducing mpt-7b: A new standard for open-source, commercially usable llms
MosaicML NLP Team. 2023 · 2023
Closest in time.
Focused transformer: Contrastive training for context scaling
Szymon Tworkowski, Konrad Staniszewski, Mikołaj Pacek, Yuhuai Wu, Henryk Michalewski, and Piotr Miłoś. 2023 · 2023
Closest in time.
DOC: Improving long story coherence with detailed outline control
Kevin Yang, Dan Klein, Nanyun Peng, and Yuandong Tian. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mirelle Candida Bueno, Carlos Gemmell, Jeff Dalton, Roberto Lotufo, and Rodrigo Nogueira. 2022 · 2022
Cited alongside, same era.
Sketching as a tool for understanding and accelerating self-attention for long sequences
Yifan Chen, Qi Zeng, Heng Ji, and Yun Yang. 2022 · 2022
Cited alongside, same era.
A length-extrapolatable transformer
Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei. 2022 · 2022
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2023 · 2023
Cited alongside, same era.
Dissecting transformer length extrapolation via the lens of receptive field analysis
Ta-Chung Chi, Ting-Han Fan, Alexander Rudnicky, and Peter Ramadge. 2023 · 2023
Cited alongside, same era.
Longnet: Scaling transformers to 1,000,000,000 tokens
Jiayu Ding, Shuming Ma, Li Dong, Xingxing Zhang, Shaohan Huang, Wenhui Wang, and Furu Wei. 2023 · 2023
Cited alongside, same era.
Boosting llm reasoning: Push the limits of few-shot learning with reinforced in-context pruning
Xijie Huang, Li Lyna Zhang, Kwang-Ting Cheng, and Mao Yang. 2023 · 2023
Cited alongside, same era.
Closest in time.
Recurrentgpt: Interactive generation of (arbitrarily) long text
Wangchunshu Zhou, Yuchen Eleanor Jiang, Peng Cui, Tiannan Wang, Zhenxin Xiao, Yifan Hou, Ryan Cotterell, and Mrinmaya Sachan. 2023 · 2023
Closest in time.
Pose: Efficient context window extension of llms via positional skip-wise training
Dawei Zhu, Nan Yang, Liang Wang, Yifan Song, Wenhao Wu, Furu Wei, and Sujian Li. 2023 · 2023
Closest in time.
Longalign: A recipe for long context alignment of large language models
Yushi Bai, Xin Lv, Jiajie Zhang, Yuze He, Ji Qi, Lei Hou, Jie Tang, Yuxiao Dong, and Juanzi Li. 2024 · 2024
Closest in time.
Longrope: Extending llm context window beyond 2 million tokens
Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang. 2024 · 2024
Closest in time.
Get more with less: Synthesizing recurrence with kv cache compression for efficient llm inference
Harry Dong, Xinyu Yang, Zhenyu Zhang, Zhangyang Wang, Yuejie Chi, and Beidi Chen. 2024 · 2024
Closest in time.
A human-inspired reading agent with gist memory of very long contexts
Kuang-Huei Lee, Xinyun Chen, Hiroki Furuta, John Canny, and Ian Fischer. 2024 · 2024
Closest in time.
Kun Luo, Zheng Liu, Shitao Xiao, and Kang Liu. 2024 · 2024
Closest in time.
Longwanjuan: Towards systematic measurement for long text quality
Kai Lv, Xiaoran Liu, Qipeng Guo, Hang Yan, Conghui He, Xipeng Qiu, and Dahua Lin. 2024 · 2024
Closest in time.
Transformers are multi-state rnns
Matanel Oren, Michael Hassid, Yossi Adi, and Roy Schwartz. 2024 · 2024
Closest in time.
Clongeval: A chinese benchmark for evaluating long-context large language models
Zexuan Qiu, Jingjing Li, Shijue Huang, Wanjun Zhong, and Irwin King. 2024 · 2024
Closest in time.
On the efficacy of eviction policy for key-value constrained generative language model inference
Siyu Ren and Kenny Q Zhu. 2024 · 2024
Closest in time.
Flexibly scaling large language models contexts through extensible tokenization
Ninglu Shao, Shitao Xiao, Zheng Liu, and Peitian Zhang. 2024 · 2024
Closest in time.
Novelqa: A benchmark for long-range novel question answering
Cunxiang Wang, Ruoxi Ning, Boqi Pan, Tonghui Wu, Qipeng Guo, Cheng Deng, Guangsheng Bao, Qian Wang, and Yue Zhang. 2024 · 2024
Closest in time.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024 · 2024
Closest in time.
Zi Yang and Nan Hua. 2024 · 2024
Closest in time.
Lv-eval: A balanced long-context benchmark with 5 length levels up to 256k
Tao Yuan, Xuefei Ning, Dong Zhou, Zhijie Yang, Shiyao Li, Minghui Zhuang, Zheyue Tan, Zhuyu Yao, Dahua Lin, Boxun Li, et al. 2024 · 2024
Closest in time.