Fetching the paper…
Reading the bibliography…
The ability of Large Language Models (LLMs) to process and generate coherent text is markedly weakened when the number of input tokens exceeds their pretraining length.
The probabilistic relevance framework: Bm25 and beyond
Robertson, S., Zaragoza, H., et al · 2009
Earlier work this paper cites.
Attention is all you need, 2017
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
The NarrativeQA reading comprehension challenge
Kočiský, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 2020
Earlier work this paper cites.
A dataset of information-seeking questions and answers anchored in research papers
Dasigi, P., Lo, K., Beltagy, I., Cohan, A., Smith, N. A., and Gardner, M · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
QMSum: A new benchmark for query-based multi-domain meeting summarization
Zhong, M., Yin, D., Yu, T., Zaidi, A., Mutuma, M., Jha, R., Awadallah, A. H., Celikyilmaz, A., Liu, Y., Qiu, X., and Radev, D · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
QuALITY: Question answering with long input texts, yes!
Pang, R. Y., Parrish, A., Joshi, N., Nangia, N., Phang, J., Chen, A., Padmakumar, V., Ma, J., Thompson, J., He, H., and Bowman, S · 2022
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation, 2022
Press, O., Smith, N. A., and Lewis, M · 2022
Earlier work this paper cites.
Roformer: Enhanced transformer with rotary position embedding, 2022
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2022
Earlier work this paper cites.
A length-extrapolatable transformer, 2022
Sun, Y., Dong, L., Patra, B., Ma, S., Huang, S., Benhaim, A., Chaudhary, V., Song, X., and Wei, F · 2022
Earlier work this paper cites.
L-eval: Instituting standardized evaluation for long context language models
An, C., Gong, S., Zhong, M., Li, M., Zhang, J., Kong, L., and Qiu, X · 2023
Earlier work this paper cites.
Introducing 100K Context Windows, 2023
Anthropic · 2023
Earlier work this paper cites.
Qwen technical report, 2023
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., Hui, B., Ji, L., Li, M., Lin, J., Lin, R., Liu, D., Liu, G., Lu, C., Lu, K., Ma, J., Men, R., Ren, X., Ren, X., Tan, C., Tan, S., Tu, J., Wang, P., Wang, S., Wang, W., Wu, S., Xu, B., Xu, J., Yang, A., Yang, H., Yang, J., Yang, S., Yao, Y., Yu, B., Yuan, H., Yuan, Z., Zhang, J., Zhang, X., Zhang, Y., Zhang, Z., Zhou, C., Zhou, J., Zhou, X., and Zhu, T · 2023
Earlier work this paper cites.
Dissecting transformer length extrapolation via the lens of receptive field analysis, 2023
Chi, T.-C., Fan, T.-H., Rudnicky, A. I., and Ramadge, P. J · 2023
Earlier work this paper cites.
Monotonic location attention for length generalization, 2023
Chowdhury, J. R. and Caragea, C · 2023
Cited alongside, same era.
Redpajama: an open dataset for training large language models, 2023
Computer, T · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning, 2023
Dao, T · 2023
Cited alongside, same era.
Lm-infinite: Simple on-the-fly length generalization for large language models, 2023
Han, C., Wang, Q., Xiong, W., Chen, Y., Ji, H., and Wang, S · 2023
Cited alongside, same era.
The impact of positional encoding on length generalization in transformers, 2023
Kazemnejad, A., Padhi, I., Ramamurthy, K. N., Das, P., and Reddy, S · 2023
Cited alongside, same era.
Prompted llms as chatbot modules for long open-domain conversation
Lee, G., Hartmann, V., Park, J., Papailiopoulos, D., and Lee, K · 2023
Rectified rotary position embeddings
Su, J · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Llama-2-7b-32k-instruct — and fine-tuning for llama-2 models with together api, 2023
Together · 2023
Later among the works it cites.
Focused transformer: Contrastive training for context scaling, 2023
Tworkowski, S., Staniszewski, K., Pacek, M., Wu, Y., Michalewski, H., and Miłoś, P · 2023
Later among the works it cites.
Leveraging large language models to power chatbots for collecting user self-reported data, 2023
Wei, J., Kim, S., Jung, H., and Kim, Y.-H · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks, 2023
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90
LMSYS · 2023
Cited alongside, same era.
Landmark attention: Random-access infinite context length for transformers
Mohtashami, A. and Jaggi, M · 2023
Cited alongside, same era.
Introducing mpt-30b: Raising the bar for open-source foundation models, 2023a
MosaicML · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Yarn: Efficient context window extension of large language models, 2023
Peng, B., Quesnelle, J., Fan, H., and Shippole, E · 2023
Cited alongside, same era.
Parallel context windows for large language models, 2023
Ratner, N., Levine, Y., Belinkov, Y., Ram, O., Magar, I., Abend, O., Karpas, E., Shashua, A., Leyton-Brown, K., and Shoham, Y · 2023
Cited alongside, same era.
Later among the works it cites.
Effective long-context scaling of foundation models
Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H · 2023
Later among the works it cites.
Compositional exemplars for in-context learning
Ye, J., Wu, Z., Feng, J., Yu, T., and Kong, L · 2023
Later among the works it cites.
Linear attention via orthogonal memory
Zhang, J., Jiang, S., Feng, J., Zheng, L., and Kong, L · 2023
Later among the works it cites.
Pose: Efficient context window extension of llms via positional skip-wise training, 2023
Zhu, D., Yang, N., Wang, L., Song, Y., Wu, W., Wei, F., and Li, S · 2023
Later among the works it cites.
Two stones hit one bird: Bilevel positional encoding for better length extrapolation, 2024
He, Z., Feng, G., Luo, S., Yang, K., He, D., Xu, J., Zhang, Z., Yang, H., and Wang, L · 2024
Closest in time.
Llm maybe longlm: Self-extend llm context window without tuning, 2024
Jin, H., Han, X., Yang, J., Jiang, Z., Liu, Z., Chang, C.-Y., Chen, H., and Hu, X · 2024
Closest in time.
Longwanjuan: Towards systematic measurement for long text quality, 2024
Lv, K., Liu, X., Guo, Q., Yan, H., He, C., Qiu, X., and Lin, D · 2024
Closest in time.
Lightning attention-2: A free lunch for handling unlimited sequence lengths in large language models
Qin, Z., Sun, W., Li, D., Shen, X., Sun, W., and Zhong, Y · 2024
Closest in time.
Learning to retrieve in-context examples for large language models, 2024
Wang, L., Yang, N., and Wei, F · 2024
Closest in time.
Infllm: Unveiling the intrinsic capacity of llms for understanding extremely long sequences with training-free memory, 2024
Xiao, C., Zhang, P., Han, X., Xiao, G., Lin, Y., Zhang, Z., Liu, Z., Han, S., and Sun, M · 2024
Closest in time.
Soaring from 4k to 400k: Extending llm’s context with activation beacon
Zhang, P., Liu, Z., Xiao, S., Shao, N., Ye, Q., and Dou, Z · 2024
Closest in time.