Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) struggle to handle long input sequences due to high memory and runtime costs.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 1911
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., and Lillicrap, T. P · 1911
Earlier work this paper cites.
Non-holographic associative memory
Willshaw, D. J., Buneman, O. P., and Longuet-Higgins, H. C · 1969
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
Hopfield, J. J · 1982
Earlier work this paper cites.
Sparse distributed memory
Kanerva, P · 1988
Earlier work this paper cites.
Retrieval and reconsolidation: toward a neurobiology of remembering
Sara, S. J · 2000
Earlier work this paper cites.
The neurobiology of consolidations, or, how stable is the engram?
Dudai, Y · 2004
Earlier work this paper cites.
Associative memory: A system-theoretical approach , volume 17
Kohonen, T · 2012
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Weston, J., Chopra, S., and Bordes, A · 2014
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Graves, A., Wayne, G., Reynolds, M., Harley, T., Danihelka, I., Grabska-Barwińska, A., Colmenarejo, S. G., Grefenstette, E., Ramalho, T., Agapiou, J., et al · 2016
Earlier work this paper cites.
Dense associative memory for pattern recognition
Krotov, D. and Hopfield, J. J · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Billion-scale similarity search with gpus, 2017
Johnson, J., Douze, M., and Jégou, H · 2017
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Stand-alone self-attention in vision models
Ramachandran, P., Parmar, N., Vaswani, A., Bello, I., Levskaya, A., and Shlens, J · 2019
Cited alongside, same era.
Analyzing the structure of attention in a transformer language model
Vig, J., Learning, M., and Belinkov, Y · 2019
Cited alongside, same era.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Cited alongside, same era.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., et al · 2022
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation, 2022
Press, O., Smith, N. A., and Lewis, M · 2022
Later among the works it cites.
Wu, Y., Rabe, M. N., Hutchins, D., and Szegedy, C · 2022
Later among the works it cites.
Adapting language models to compress contexts
Chevalier, A., Wettig, A., Ajith, A., and Chen, D · 2023
Later among the works it cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Cited alongside, same era.
Generative language modeling for automated theorem proving
Polu, S. and Sutskever, I · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Cited alongside, same era.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Cited alongside, same era.
Hopfield networks is all you need
Ramsauer, H., Schäfl, B., Lehner, J., Seidl, P., Widrich, M., Gruber, L., Holzleitner, M., Adler, T., Kreil, D., Kopp, M. K., et al · 2021
Cited alongside, same era.
Biological learning in key-value memory networks
Tyulmankov, D., Fang, C., Vadaparty, A., and Yang, G. R · 2021
Cited alongside, same era.
Later among the works it cites.
Longnet: Scaling transformers to 1,000,000,000 tokens
Ding, J., Ma, S., Dong, L., Zhang, X., Huang, S., Wang, W., Zheng, N., and Wei, F · 2023
Later among the works it cites.
In-context autoencoder for context compression in a large language model
Ge, T., Hu, J., Wang, X., Chen, S.-Q., and Wei, F · 2023
Later among the works it cites.
Medeval: A multi-level, multi-task, and multi-domain medical benchmark for language model evaluation
He, Z., Wang, Y., Yan, A., Liu, Y., Chang, E., Gentili, A., McAuley, J., and Hsu, C.-n · 2023
Later among the works it cites.
Lost in the middle: How language models use long contexts
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P · 2023
Later among the works it cites.
Learning to compress prompts with gist tokens
Mu, J., Li, X. L., and Goodman, N · 2023
Later among the works it cites.
Memgpt: Towards llms as operating systems
Packer, C., Fang, V., Patil, S. G., Lin, K., Wooders, S., and Gonzalez, J. E · 2023
Later among the works it cites.
Memoria: Hebbian memory architecture for human-like sequential processing
Park, S. and Bak, J · 2023
Later among the works it cites.
Focused transformer: Contrastive training for context scaling
Tworkowski, S., Staniszewski, K., Pacek, M., Wu, Y., Michalewski, H., and Miłoś, P · 2023
Later among the works it cites.
Augmenting language models with long-term memory
Wang, W., Dong, L., Cheng, H., Liu, X., Yan, X., Gao, J., and Wei, F · 2023
Later among the works it cites.