Fetching the paper…
Reading the bibliography…
Equipping large language models (LLMs) with latent-space memory has attracted increasing attention as they can extend the context window of existing language models.
What does bert learn about the structure of language?
Jawahar, G., Sagot, B., and Seddah, D · 2019
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Memory transformer with hierarchical attention for long document processing
Al Adel, A. and Burtsev, M. S · 2021
Earlier work this paper cites.
Nncp v2: Lossless data compression with transformer
Bellard, F · 2021
Earlier work this paper cites.
How many layers and why? an analysis of the model depth in transformers
Simoulin, A. and Crabbé, B · 2021
Earlier work this paper cites.
Recurrent memory transformer
Bulatov, A., Kuratov, Y., and Burtsev, M. S · 2022
Earlier work this paper cites.
Memorizing transformers
Wu, Y., Rabe, M. N., Hutchins, D., and Szegedy, C · 2022
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., et al · 2023
Earlier work this paper cites.
Scaling transformer to 1m tokens and beyond with rmt
Bulatov, A., Kuratov, Y., Kapushev, Y., and Burtsev, M. S · 2023
Earlier work this paper cites.
Chatdb: Augmenting llms with databases as their symbolic memory
Hu, C., Fu, J., Du, C., Luo, S., Zhao, J., and Zhao, H · 2023
Earlier work this paper cites.
Memgpt: Towards llms as operating systems
Packer, C., Fang, V., Patil, S. G., Lin, K., Wooders, S., and Gonzalez, J. E · 2023
Earlier work this paper cites.
Augmenting language models with long-term memory
Wang, W., Dong, L., Cheng, H., Liu, X., Yan, X., Gao, J., and Wei, F · 2023
Cited alongside, same era.
H2O: heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C. W., Wang, Z., and Chen, B · 2023
Cited alongside, same era.
Memorybank: Enhancing large language models with long-term memory
Zhong, W., Guo, L., Gao, Q., and Wang, Y · 2023
Cited alongside, same era.
Recurrentgpt: Interactive generation of (arbitrarily) long text
Zhou, W., Jiang, Y. E., Cui, P., Wang, T., Xiao, Z., Hou, Y., Cotterell, R., and Sachan, M · 2023
Cited alongside, same era.
Training large language models to reason in a continuous latent space
Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J., and Tian, Y · 2024
Later among the works it cites.
Camelot: Towards large language models with training-free consolidated associative memory
He, Z., Karlinsky, L., Kim, D., McAuley, J., Krotov, D., and Feris, R · 2024
Later among the works it cites.
Snapkv: LLM knows what you are looking for before generation
Li, Y., Huang, Y., Yang, B., Venkitesh, B., Locatelli, A., Ye, H., Cai, T., Lewis, P., and Chen, D · 2024
Later among the works it cites.
Memllm: Finetuning llms to use an explicit read-write memory
Modarressi, A., Köksal, A., Imani, A., Fayyaz, M., and Schütze, H · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Belcak, P. and Wattenhofer, R · 2024
Cited alongside, same era.
MinPrompt: Graph-based minimal prompt data augmentation for few-shot question answering
Chen, X., Jiang, J.-Y., Chang, W.-C., Hsieh, C.-J., Yu, H.-F., and Wang, W · 2024
Cited alongside, same era.
Larimar: Large language models with episodic memory control
Das, P., Chaudhury, S., Nelson, E., Melnyk, I., Swaminathan, S., Dai, S., Lozano, A. C., Kollias, G., Chenthamarakshan, V., Navrátil, J., Dan, S., and Chen, P · 2024
Cited alongside, same era.
The llama 3 herd of models
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., Rao, A., Zhang, A., Rodriguez, A., Gregerson, A., Spataru, A., Rozière, B., Biron, B., Tang, B., Chern, B., Caucheteux, C., Nayak, C., Bi, C., Marra, C., McConnell, C., Keller, C., Touret, C., Wu, C., Wong, C., Ferrer, C. C., Nikolaidis, C., Allonsius, D., Song, D., Pintz, D., Livshits, D., Esiobu, D., Choudhary, D., Mahajan, D., Garcia-Olano, D., Perino, D., Hupkes, D., Lakomkin, E., AlBadawy, E., Lobanova, E., Dinan, E., Smith, E. M., Radenovic, F., Zhang, F., Synnaeve, G., Lee, G., Anderson, G. L., Nail, G., Mialon, G., Pang, G., Cucurell, G., Nguyen, H., Korevaar, H., Xu, H., Touvron, H., Zarov, I., Ibarra, I. A., Kloumann, I. M., Misra, I., Evtimov, I., Copet, J., Lee, J., Geffert, J., Vranes, J., Park, J., Mahadeokar, J., Shah, J., van der Linde, J., Billock, J., Hong, J., Lee, J., Fu, J., Chi, J., Huang, J., Liu, J., Wang, J., Yu, J., Bitton, J., Spisak, J., Park, J., Rocca, J., Johnstun, J., Saxe, J., Jia, J., Alwala, K. V., Upasani, K., Plawiak, K., Li, K., Heafield, K., Stone, K., and et al · 2024
Cited alongside, same era.
Language is primarily a tool for communication rather than thought
Fedorenko, E., Piantadosi, S. T., and Gibson, E. A · 2024
Cited alongside, same era.
Data engineering for scaling language models to 128k context
Fu, Y., Panda, R., Niu, X., Yue, X., Hajishirzi, H., Kim, Y., and Peng, H · 2024
Cited alongside, same era.
In-context autoencoder for context compression in a large language model
Ge, T., Hu, J., Wang, L., Wang, X., Chen, S., and Wei, F · 2024
Cited alongside, same era.
Hipporag: Neurobiologically inspired long-term memory for large language models
Gutiérrez, B. J., Shu, Y., Gu, Y., Yasunaga, M., and Su, Y · 2024
Cited alongside, same era.
Park, S. and Bak, J · 2024
Later among the works it cites.
The fineweb datasets: Decanting the web for the finest text data at scale
Penedo, G., Kydlícek, H., Allal, L. B., Lozhkov, A., Mitchell, M., Raffel, C., von Werra, L., and Wolf, T · 2024
Later among the works it cites.
Who’s who: Large language models meet knowledge conflicts in practice
Pham, Q. H., Ngo, H., Luu, A. T., and Nguyen, D. Q · 2024
Later among the works it cites.
An enhanced text compression approach using transformer-based language models
Rahman, C. M., Sobhani, M. E., Rodela, A. T., and Shatabda, S · 2024
Later among the works it cites.
Memoryprompt: A light wrapper to improve context tracking in pre-trained language models
Rakotonirina, N. C. and Baroni, M · 2024
Later among the works it cites.
Memory 3 {}^{\mbox{3}} : Language modeling with explicit memory
Yang, H., Lin, Z., Wang, W., Wu, H., Li, Z., Tang, B., Wei, W., Wang, J., Tang, Z., Song, S., Xi, C., Yu, Y., Chen, K., Xiong, F., Tang, L., and E, W · 2024
Later among the works it cites.
Explicit memory learning with expectation maximization
Yin, Z., Sun, Q., Guo, Q., Zeng, Z., Cheng, Q., Qiu, X., and Huang, X · 2024
Later among the works it cites.
∞ \infty bench: Extending long context evaluation beyond 100k tokens
Zhang, X., Chen, Y., Hu, S., Xu, Z., Chen, J., Hao, M., Han, X., Thai, Z., Wang, S., Liu, Z., et al · 2024
Later among the works it cites.