Fetching the paper…
Reading the bibliography…
Extrapolating ultra-long contexts (text length >128K) remains a major challenge for large language models (LLMs), as most training-free extrapolation methods are not only severely limited by memory bottlenecks, but also suffer from the attention sink, which restricts their scalability and effectiveness in practice.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J · 2018
Earlier work this paper cites.
The narrativeqa reading comprehension challenge
Kočiskỳ, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E · 2018
Earlier work this paper cites.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Dong, Y., Cordonnier, J.-B., and Loukas, A · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M · 2021
Earlier work this paper cites.
Kerple: Kernelized relative positional embedding for length extrapolation
Chi, T.-C., Fan, T.-H., Ramadge, P. J., and Rudnicky, A · 2022
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation, 2022
Press, O., Smith, N. A., and Lewis, M · 2022
Earlier work this paper cites.
Parallel context windows for large language models
Ratner, N., Levine, Y., Belinkov, Y., Ram, O., Magar, I., Abend, O., Karpas, E., Shashua, A., Leyton-Brown, K., and Shoham, Y · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., et al · 2023
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Earlier work this paper cites.
Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression
Jiang, H., Wu, Q., Luo, X., Li, D., Lin, C.-Y., Yang, Y., and Qiu, L · 2023
Earlier work this paper cites.
Yarn: Efficient context window extension of large language models
Peng, B., Quesnelle, J., Fan, H., and Shippole, E · 2023
Earlier work this paper cites.
Attention sorting combats recency bias in long context language models
Peysakhovich, A. and Lerer, A · 2023
Earlier work this paper cites.
Rectified rotary position embeddings
Su, J · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Efficient streaming language models with attention sinks, 2023
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Cited alongside, same era.
Dq-lore: Dual queries with low rank approximation re-ranking for in-context learning
Xiong, J., Li, Z., Zheng, C., Guo, Z., Yin, Y., Xie, E., Yang, Z., Cao, Q., Wang, H., Han, X., et al · 2023
Cited alongside, same era.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., et al · 2023
Cited alongside, same era.
Lost in the middle: How language models use long contexts
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Later among the works it cites.
D2o: Dynamic discriminative operations for efficient generative inference of large language models
Wan, Z., Wu, X., Zhang, Y., Xin, Y., Tao, C., Zhu, Z., Wang, X., Luo, S., Xiong, J., and Zhang, M · 2024
Later among the works it cites.
Infllm: Training-free long-context extrapolation for llms with an efficient context memory
Xiao, C., Zhang, P., Han, X., Xiao, G., Lin, Y., Zhang, Z., Liu, Z., and Sun, M · 2024
Later among the works it cites.
Uncomp: Uncertainty-aware long-context compressor for efficient large language model inference
Xiong, J., Shen, J., Ye, F., Tao, C., Wan, Z., Lu, J., Wu, X., Zheng, C., Guo, Z., Kong, L., et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Acharya, S., Jia, F., and Ginsburg, B · 2024
Cited alongside, same era.
Llama 3: A family of large language models
AI, M · 2024
Cited alongside, same era.
Training-free long-context scaling of large language models
An, C., Huang, F., Zhang, J., Gong, S., Qiu, X., Zhou, C., and Kong, L · 2024
Cited alongside, same era.
Chen, Y., Lv, A., Luan, J., Wang, B., and Liu, W · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Data engineering for scaling language models to 128k context
Fu, Y., Panda, R., Niu, X., Yue, X., Hajishirzi, H., Kim, Y., and Peng, H · 2024
Cited alongside, same era.
Chatglm: A family of large language models from glm-130b to glm-4 all tools
GLM, T., Zeng, A., Xu, B., Wang, B., Zhang, C., Yin, D., Zhang, D., Rojas, D., Feng, G., Zhao, H., et al · 2024
Cited alongside, same era.
Contextual position encoding: Learning to count what’s important
Golovneva, O., Wang, T., Weston, J., and Sukhbaatar, S · 2024
Cited alongside, same era.
Later among the works it cites.
Chatqa 2: Bridging the gap to proprietary llms in long context and rag capabilities
Xu, P., Ping, W., Wu, X., Xu, C., Liu, Z., Shoeybi, M., and Catanzaro, B · 2024
Later among the works it cites.
Yu, Z., Wang, Z., Fu, Y., Shi, H., Shaikh, K., and Lin, Y. C · 2024
Later among the works it cites.
infinite bench: Extending long context evaluation beyond 100k tokens
Zhang, X., Chen, Y., Hu, S., Xu, Z., Chen, J., Hao, M., Han, X., Thai, Z., Wang, S., Liu, Z., et al · 2024
Later among the works it cites.
Dape: Data-adaptive positional encoding for length extrapolation
Zheng, C., Gao, Y., Shi, H., Huang, M., Li, J., Xiong, J., Ren, X., Ng, M., Jiang, X., Li, Z., et al · 2024
Later among the works it cites.
Accelerating inference of retrieval-augmented generation via sparse context selection
Zhu, Y., Gu, J.-C., Sikora, C., Ko, H., Liu, Y., Lin, C.-C., Shu, L., Luo, L., Meng, L., Liu, B., et al · 2024
Later among the works it cites.
Yi-34b-200k, 2023a
01.AI · 2025
Closest in time.
Yi-6b-200k, 2023b
01.AI · 2025
Closest in time.
Kimi chat, 2023
AI, M · 2025
Closest in time.
Model card and evaluations for claude models, 2023
Anthropic · 2025
Closest in time.
Add ntk-aware interpolation ”by parts” correction
bloc97 · 2025
Closest in time.