Fetching the paper…
Reading the bibliography…
LongRoPE2 is a novel approach that extends the effective context window of pre-trained large language models (LLMs) to the target length, while preserving the performance on the original shorter context window.
Generating long sequences with sparse transformers, 2019
Child, R., Gray, S., Radford, A., and Sutskever, I · 1904
Earlier work this paper cites.
Longformer: The long-document transformer, 2020
Beltagy, I., Peters, M. E., and Cohan, A · 2004
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention, 2020
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2006
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Fourier features let networks learn high frequency functions in low dimensional domains
Tancik, M., Srinivasan, P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Singhal, U., Ramamoorthi, R., Barron, J., and Ng, R · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A · 2020
Earlier work this paper cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2021
Earlier work this paper cites.
Longt5: Efficient text-to-text transformer for long sequences, 2022
Guo, M., Ainslie, J., Uthus, D., Ontanon, S., Ni, J., Sung, Y.-H., and Yang, Y · 2022
Earlier work this paper cites.
Gpt-4 technical report, 2023
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., et al · 2023
Earlier work this paper cites.
Extending context window of large language models via positional interpolation
Chen, S., Wong, S., Chen, L., and Tian, Y · 2023
Earlier work this paper cites.
Redpajama: An open source recipe to reproduce llama training dataset, 2023
Computer, T · 2023
Earlier work this paper cites.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Earlier work this paper cites.
Longnet: Scaling transformers to 1,000,000,000 tokens, 2023
Ding, J., Ma, S., Dong, L., Zhang, X., Huang, S., Wang, W., Zheng, N., and Wei, F · 2023
Earlier work this paper cites.
Lm-infinite: Simple on-the-fly length generalization for large language models
Han, C., Wang, Q., Xiong, W., Chen, Y., Ji, H., and Wang, S · 2023
Earlier work this paper cites.
Needle in a haystack - pressure testing llms, 2023
Kamradt, G · 2023
Cited alongside, same era.
Starcoder: may the source be with you!
Li, R., Allal, L. B., Zi, Y., Muennighoff, N., Kocetkov, D., Mou, C., Marone, M., Akiki, C., Li, J., Chim, J., et al · 2023
Cited alongside, same era.
Scaling laws of rope-based extrapolation
Liu, X., Yan, H., Zhang, S., An, C., Qiu, X., and Lin, D · 2023
Cited alongside, same era.
Ntk-aware scaled rope allows llama models to have extended (8k+) context size without any fine-tuning and minimal perplexity degration, 2023
LocalLLaMA · 2023
Cited alongside, same era.
Yarn: Efficient context window extension of large language models
Peng, B., Quesnelle, J., Fan, H., and Shippole, E · 2023
A real-world webagent with planning, long context understanding, and program synthesis, 2024
Gur, I., Furuta, H., Huang, A., Safdari, M., Matsuo, Y., Eck, D., and Faust, A · 2024
Later among the works it cites.
Ruler: What’s the real context size of your long-context language models?
Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., Zhang, Y., and Ginsburg, B · 2024
Later among the works it cites.
Jeong, S., Baek, J., Cho, S., Hwang, S. J., and Park, J. C · 2024
Later among the works it cites.
Jamba: A hybrid transformer-mamba language model, 2024
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., Abend, O., Alon, R., Asida, T., Bergman, A., Glozman, R., Gokhman, M., Manevich, A., Ratner, N., Rozen, N., Shwartz, E., Zusman, M., and Shoham, Y · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone, 2024
Abdin, M., Aneja, J., Awadalla, H., Awadallah, A., Awan, A. A., Bach, N., Bahree, A., Bakhtiari, A., Bao, J., Behl, H., Benhaim, A., Bilenko, M., Bjorck, J., Bubeck, S., Cai, M., Cai, Q., Chaudhary, V., Chen, D., Chen, D., Chen, W., Chen, Y.-C., Chen, Y.-L., Cheng, H., Chopra, P., Dai, X., Dixon, M., Eldan, R., Fragoso, V., Gao, J., Gao, M., Gao, M., Garg, A., Giorno, A. D., Goswami, A., Gunasekar, S., Haider, E., Hao, J., Hewett, R. J., Hu, W., Huynh, J., Iter, D., Jacobs, S. A., Javaheripi, M., Jin, X., Karampatziakis, N., Kauffmann, P., Khademi, M., Kim, D., Kim, Y. J., Kurilenko, L., Lee, J. R., Lee, Y. T., Li, Y., Li, Y., Liang, C., Liden, L., Lin, X., Lin, Z., Liu, C., Liu, L., Liu, M., Liu, W., Liu, X., Luo, C., Madan, P., Mahmoudzadeh, A., Majercak, D., Mazzola, M., Mendes, C. C. T., Mitra, A., Modi, H., Nguyen, A., Norick, B., Patra, B., Perez-Becker, D., Portet, T., Pryzant, R., Qin, H., Radmilac, M., Ren, L., de Rosa, G., Rosset, C., Roy, S., Ruwase, O., Saarikivi, O., Saied, A., Salim, A., Santacroce, M., Shah, S., Shang, N., Sharma, H., Shen, Y., Shukla, S., Song, X., Tanaka, M., Tupini, A., Vaddamanu, P., Wang, C., Wang, G., Wang, L., Wang, S., Wang, X., Wang, Y., Ward, R., Wen, W., Witte, P., Wu, H., Wu, X., Wyatt, M., Xiao, B., Xu, C., Xu, J., Xu, W., Xue, J., Yadav, S., Yang, F., Yang, J., Yang, Y., Yang, Z., Yu, D., Yuan, L., Zhang, C., Zhang, C., Zhang, J., Zhang, L. L., Zhang, Y., Zhang, Y., Zhang, Y., and Zhou, X · 2024
Cited alongside, same era.
Training-free long-context scaling of large language models, 2024
An, C., Huang, F., Zhang, J., Gong, S., Qiu, X., Zhou, C., and Kong, L · 2024
Cited alongside, same era.
Rq-rag: Learning to refine queries for retrieval augmented generation, 2024
Chan, C.-M., Xu, C., Yuan, R., Luo, H., Xue, W., Guo, Y., and Fu, J · 2024
Cited alongside, same era.
Longrope: Extending llm context window beyond 2 million tokens
Ding, Y., Zhang, L. L., Zhang, C., Xu, Y., Shang, N., Xu, J., Yang, F., and Yang, M · 2024
Cited alongside, same era.
Multi-view content-aware indexing for long document retrieval, 2024
Dong, K., Deik, D. G. X., Lee, Y. Q., Zhang, H., Li, X., Zhang, C., and Liu, Y · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
What is wrong with perplexity for long-context language modeling?
Fang, L., Wang, Y., Liu, Z., Zhang, C., Jegelka, S., Gao, J., Ding, B., and Wang, Y · 2024
Cited alongside, same era.
nnscaler: Constraint-guided parallelization plan generation for deep learning training
Lin, Z., Miao, Y., Zhang, Q., Yang, F., Zhu, Y., Li, C., Maleki, S., Cao, X., Shang, N., Yang, Y., Xu, W., Yang, M., Zhang, L., and Zhou, L · 2024
Later among the works it cites.
Fineweb-edu: the finest collection of educational content, 2024
Lozhkov, A., Ben Allal, L., von Werra, L., and Wolf, T · 2024
Later among the works it cites.
Luo, K., Liu, Z., Xiao, S., and Liu, K · 2024
Later among the works it cites.
Llama3.2: Revolutionizing edge ai and vision with open, customizable models, 2024
Meta · 2024
Later among the works it cites.
Samba: Simple hybrid state space models for efficient unlimited context language modeling, 2024
Ren, L., Liu, Y., Lu, Y., Shen, Y., Liang, C., and Chen, W · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models, September 2024
Team, Q · 2024
Later among the works it cites.
When precision meets position: Bfloat16 breaks down rope in long-context training
Wang, H., Liu, Q., Du, C., Zhu, T., Du, C., Kawaguchi, K., and Pang, T · 2024
Later among the works it cites.
Redpajama: an open dataset for training large language models
Weber, M., Fu, D. Y., Anthony, Q., Oren, Y., Adams, S., Alexandrov, A., Lyu, X., Nguyen, H., Yao, X., Adams, V., Athiwaratkun, B., Chalamala, R., Chen, K., Ryabinin, M., Dao, T., Liang, P., Ré, C., Rish, I., and Zhang, C · 2024
Later among the works it cites.
Robustifying state-space models for long sequences via approximate diagonalization
Yu, A., Nigmetov, A., Morozov, D., Mahoney, M. W., and Erichson, N. B · 2024
Later among the works it cites.
Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence
Zhu, Q., Guo, D., Shao, Z., Yang, D., Wang, P., Xu, R., Wu, Y., Li, Y., Gao, H., Ma, S., et al · 2024
Later among the works it cites.
Hipporag: Neurobiologically inspired long-term memory for large language models, 2025
Gutiérrez, B. J., Shu, Y., Gu, Y., Yasunaga, M., and Su, Y · 2025
Closest in time.