Fetching the paper…
Reading the bibliography…
Current large language models (LLMs) often perform poorly on simple fact retrieval tasks.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan Salakhutdinov. 2019 · 1901
Earlier work this paper cites.
Representation of associated data by matrix operators
Teuvo Kohonen and Matti Ruohonen. 1973 · 1973
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Mikhail S. Burtsev, Yuri Kuratov, Anton Peganov, and Grigory V. Sapunov. 2021 · 2006
Earlier work this paper cites.
Rewriting a deep generative model
David Bau, Steven Liu, Tongzhou Wang, Jun-Yan Zhu, and Antonio Torralba. 2020 · 2007
Earlier work this paper cites.
The kanerva machine: A generative distributed memory
Yan Wu, Greg Wayne, Alex Graves, and Timothy Lillicrap. 2018 · 2018
Earlier work this paper cites.
Optimus: Organizing sentences via pre-trained modeling of a latent space
Chunyuan Li, Xiang Gao, Yuan Li, Baolin Peng, Xiujun Li, Yizhe Zhang, and Jianfeng Gao. 2020 · 2020
Cited alongside, same era.
Generative pseudo-inverse memory
Kha Pham, Hung Le, Man Ngo, Truyen Tran, Bao Ho, and Svetha Venkatesh. 2021 · 2021
Cited alongside, same era.
Recurrent memory transformer
Aydar Bulatov, Yury Kuratov, and Mikhail Burtsev. 2022 · 2022
Cited alongside, same era.
Needle In A Haystack - pressure testing LLMs
Gregory Kamradt. 2023 · 2023
Cited alongside, same era.
Landmark attention: Random-access infinite context length for transformers
Amirkeivan Mohtashami and Martin Jaggi. 2023 · 2023
Cited alongside, same era.
Flexgen: high-throughput generative inference of large language models with a single gpu
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023 · 2023
Later among the works it cites.
Larimar: Large language models with episodic memory control
Payel Das, Subhajit Chaudhury, Elliot Nelson, Igor Melnyk, Sarath Swaminathan, Sihui Dai, Aurélie Lozano, Georgios Kollias, Vijil Chenthamarakshan, Jiří, Navrátil, Soham Dan, and Pin-Yu Chen. 2024 · 2024
Closest in time.
Transformerfam: Feedback attention is working memory
Dongseong Hwang, Weiran Wang, Zhuoyuan Huo, Khe Chai Sim, and Pedro Moreno Mengibar. 2024 · 2024
Closest in time.
In search of needles in a 11m haystack: Recurrent memory finds what llms miss
Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Dmitry Sorokin, Artyom Sorokin, and Mikhail Burtsev. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tsendsuren Munkhdalai, Manaal Faruqui, and Siddharth Gopal. 2024 · 2024
Closest in time.