Fetching the paper…
Reading the bibliography…
Long-context modeling presents a significant challenge for transformer-based large language models (LLMs) due to the quadratic complexity of the self-attention mechanism and issues with length extrapolation caused by pretraining exclusively on short inputs.
The development of strategic readers
Paris, S. G., Wasik, B., and Turner, J. C · 1991
Earlier work this paper cites.
The art of computer programming , volume 3
Knuth, D. E · 1997
Earlier work this paper cites.
Verbal protocols of reading: The nature of constructively responsive reading
Pressley, M. and Afflerbach, P · 2012
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Confidence modeling for neural semantic parsing
Dong, L., Quirk, C., and Lapata, M · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Rethinking attention with performers
Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., et al · 2020
Earlier work this paper cites.
Cogltx: Applying bert to long texts
Ding, M., Zhou, C., Yang, H., and Tang, J · 2020
Earlier work this paper cites.
Selective question answering under domain shift
Kamath, A., Jia, R., and Liang, P · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Random feature attention
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N., and Kong, L · 2020
Earlier work this paper cites.
Blockwise self-attention for long document understanding
Qiu, J., Ma, H., Levy, O., Yih, W.-t., Wang, S., and Tang, J · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
How can we know when language models know? on the calibration of language models for question answering
Jiang, Z., Araki, J., Ding, H., and Neubig, G · 2021
Earlier work this paper cites.
Luna: Linear unified nested attention
Ma, X., Kong, X., Wang, S., Zhou, C., May, J., Ma, H., and Zettlemoyer, L · 2021
Cited alongside, same era.
Narrative question answering with cutting-edge open-domain qa techniques: A comprehensive study
Mou, X., Yang, C., Yu, M., Yao, B., Guo, X., Potdar, S., and Su, H · 2021
Cited alongside, same era.
Memorizing transformers
Wu, Y., Rabe, M. N., Hutchins, D., and Szegedy, C · 2021
Cited alongside, same era.
Longt5: Efficient text-to-text transformer for long sequences
Guo, M., Ainslie, J., Uthus, D. C., Ontanon, S., Ni, J., Sung, Y.-H., and Yang, Y · 2022
Cited alongside, same era.
Leveraging locality in abstractive text summarization
Liu, Y., Ni, A., Nan, L., Deb, B., Zhu, C., Hassan, A., and Radev, D · 2022
Cited alongside, same era.
Fourierformer: Transformer meets generalized fourier integral theorem
Nguyen, T., Pham, M., Nguyen, T., Nguyen, K., Osher, S., and Ho, N · 2022
Deja vu: Contextual sparsity for efficient LLMs at inference time
Liu, Z., Wang, J., Dao, T., Zhou, T., Yuan, B., Song, Z., Shrivastava, A., Zhang, C., Tian, Y., Re, C., et al · 2023
Later among the works it cites.
Landmark attention: Random-access infinite context length for transformers
Mohtashami, A. and Jaggi, M · 2023
Later among the works it cites.
Yarn: Efficient context window extension of large language models
Peng, B., Quesnelle, J., Fan, H., and Shippole, E · 2023
Later among the works it cites.
Replug: Retrieval-augmented black-box language models
Shi, W., Min, S., Yasunaga, M., Seo, M., James, R., Lewis, M., Zettlemoyer, L., and Yih, W.-t · 2023
Later among the works it cites.
Effective long-context scaling of foundation models
Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Summn: A multi-stage summarization framework for long input dialogues and documents
Zhang, Y., Ni, A., Mao, Z., Wu, C. H., Zhu, C., Deb, B., Awadallah, A., Radev, D., and Zhang, R · 2022
Cited alongside, same era.
Walking down the memory maze: Beyond context limit through interactive reading
Chen, H., Pasunuru, R., Weston, J., and Celikyilmaz, A · 2023
Cited alongside, same era.
Monotonic location attention for length generalization
Chowdhury, J. R. and Caragea, C · 2023
Cited alongside, same era.
A survey on long text modeling with transformers
Dong, Z., Tang, T., Li, L., and Zhao, W. X · 2023
Cited alongside, same era.
Simple hardware-efficient long convolutions for sequence modeling
Fu, D. Y., Epstein, E. L., Nguyen, E., Thomas, A. W., Zhang, M., Dao, T., Rudra, A., and Re, C · 2023
Cited alongside, same era.
Lm-infinite: Simple on-the-fly length generalization for large language models
Han, C., Wang, Q., Xiong, W., Chen, Y., Ji, H., and Wang, S · 2023
Cited alongside, same era.
Later among the works it cites.
Can you follow me? testing situational understanding for ChatGPT
Yang, C. and Ettinger, A · 2023
Later among the works it cites.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al · 2023
Later among the works it cites.
Mind2web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2024
Closest in time.
Data engineering for scaling language models to 128k context
Fu, Y., Panda, R., Niu, X., Yue, X., Hajishirzi, H., Kim, Y., and Peng, H · 2024
Closest in time.
Random-access infinite context length for transformers
Mohtashami, A. and Jaggi, M · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Closest in time.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2024
Closest in time.
Yang, Z. and Hua, N · 2024
Closest in time.