Fetching the paper…
Reading the bibliography…
Existing Large Language Models (LLMs) usually remain static after deployment, which might make it hard to inject new knowledge into the model.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 1904
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2004
Earlier work this paper cites.
Burtsev, M. S. and Sapunov, G. V · 2006
Earlier work this paper cites.
Weston, J., Chopra, S., and Bordes, A · 2014
Earlier work this paper cites.
End-to-end memory networks
Sukhbaatar, S., Weston, J., Fergus, R., et al · 2015
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
Levy, O., Seo, M., Choi, E., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 2019
Earlier work this paper cites.
Meshed-memory transformer for image captioning
Cornia, M., Stefanini, M., Baraldi, L., and Cucchiara, R · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Earlier work this paper cites.
Modifying memories in transformer models
Zhu, C., Rawat, A. S., Zaheer, M., Bhojanapalli, S., Li, D., Yu, F. X., and Kumar, S · 2020
Earlier work this paper cites.
Editing factual knowledge in language models
Cao, N. D., Aziz, W., and Titov, I · 2021
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M · 2021
Cited alongside, same era.
Global memory transformer for processing long documents
Adel, A. A · 2022
Cited alongside, same era.
Recurrent memory transformer
Bulatov, A., Kuratov, Y., and Burtsev, M. S · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Cited alongside, same era.
Fast model editing at scale
Mitchell, E., Lin, C., Bosselut, A., Finn, C., and Manning, C. D · 2022
Ret-llm: Towards a general read-write memory for large language models
Modarressi, A., Imani, A., Fayyaz, M., and Schütze, H · 2023
Later among the works it cites.
Efficient memory-enhanced transformer for long-document summarization in low-resource regimes
Moro, G., Ragazzi, L., Valgimigli, L., Frisoni, G., Sartori, C., and Marfia, G · 2023
Later among the works it cites.
A length-extrapolatable transformer
Sun, Y., Dong, L., Patra, B., Ma, S., Huang, S., Benhaim, A., Chaudhary, V., Song, X., and Wei, F · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Canton-Ferrer, C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Memformer: A memory-augmented transformer for sequence modeling
Wu, Q., Lan, Z., Qian, K., Gu, J., Geramifard, A., and Yu, Z · 2022
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., et al · 2023
Cited alongside, same era.
Redpajama: an open dataset for training large language models, 2023
Computer, T · 2023
Cited alongside, same era.
Openllama: An open reproduction of llama, May 2023
Geng, X. and Liu, H · 2023
Cited alongside, same era.
Longnet: Scaling transformers to 1,000,000,000 tokens
Jiayu, D., Shuming, M., Li, D., Xingxing, Z., Shaohan, H., Wenhui, W., and Wei†, F · 2023
Cited alongside, same era.
Extending context window of large language models via positional interpolation
Chen, S., Wong, S., Chen, L., and Tian, Y
Cited in the paper.
Later among the works it cites.
Focused transformer: Contrastive training for context scaling
Tworkowski, S., Staniszewski, K., Pacek, M., Wu, Y., Michalewski, H., and Miłoś, P · 2023
Later among the works it cites.
Augmenting language models with long-term memory
Wang, W., Dong, L., Cheng, H., Liu, X., Yan, X., Gao, J., and Wei, F · 2023
Later among the works it cites.
Effective long-context scaling of foundation models
Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., et al · 2023
Later among the works it cites.
Editing large language models: Problems, methods, and opportunities
Yao, Y., Wang, P., Tian, B., Cheng, S., Li, Z., Deng, S., Chen, H., and Zhang, N · 2023
Later among the works it cites.
Can we edit factual knowledge by in-context learning?
Zheng, C., Li, L., Dong, Q., Fan, Y., Wu, Z., Xu, J., and Chang, B · 2023
Later among the works it cites.
Memorybank: Enhancing large language models with long-term memory
Zhong, W., Guo, L., Gao, Q., and Wang, Y · 2023
Later among the works it cites.
Unimem: Towards a unified view of long-context large language models
Fang, J., Tang, L., Bi, H., Qin, Y., Sun, S., Li, Z., Li, H., Li, Y., Cong, X., Yan, Y., et al · 2024
Closest in time.