Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) represent a landmark achievement in Artificial Intelligence (AI), demonstrating unprecedented proficiency in procedural tasks such as text generation, code completion, and conversational coherence.
On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples
W. R. Thompson. 1933 · 1933
Earlier work this paper cites.
Episodic and semantic memory
E. Tulving. 1972 · 1972
Earlier work this paper cites.
ACT-R: A simple theory of complex cognition
J. R. Anderson and C. Lebiere. 1996 · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction . Vol. 1
R. S. Sutton, A. G. Barto, et al · 1998
Earlier work this paper cites.
The role of the basal ganglia in learning and memory: Insight from Parkinson’s disease
K. Foerde and D. Shohamy. 2011 · 2011
Earlier work this paper cites.
A large-scale model of the functioning brain
C. Eliasmith, T. C. Stewart, X. Choo, T. Bekolay, T. DeWolf, Y. Tang, and D. Rasmussen. 2012 · 2012
Earlier work this paper cites.
Two cortical systems for memory-guided behaviour
C. Ranganath and M. Ritchey. 2012 · 2012
Earlier work this paper cites.
Sparse and distributed coding of episodic memory in neurons of the human hippocampus
J. T. Wixted, L. R. Squire, Y. Jang, M. H. Papesh, S. D. Goldinger, J. R. Kuhn, K. A. Smith, D. M. Treiman, and P. N. Steinmetz. 2014 · 2014
Earlier work this paper cites.
An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks
I. J. Goodfellow, M. Mirza, D. Xiao, A. Courville, and Y. Bengio. 2015 · 2015
Earlier work this paper cites.
The Two Settings of Kind and Wicked Learning Environments
R. M. Hogarth, T. Lejarraga, and E. Soyer. 2015 · 2015
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, D. Hassabis, C. Clopath, D. Kumaran, and R. Hadsell. 2017 · 2017
Cited alongside, same era.
The Role of Hippocampal Replay in Memory and Planning
H. F. Ólafsdóttir, D. Bush, and C. Barry. 2018 · 2017
Cited alongside, same era.
Attention is All you Need. In Advances in Neural Information Processing Systems (NeurIPS ’17, Vol. 30) . Curran Associates, Inc
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. 2017 · 2017
Cited alongside, same era.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems , Vol. 33. Curran Associates, Inc., 9459–9474
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela. 2020 · 2020
Cited alongside, same era.
Energy efficient synaptic plasticity
H. L. Li and M. CW van Rossum. 2020 · 2020
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems , Vol. 35. Curran Associates, Inc., 24824–24837
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou. 2022 · 2022
Later among the works it cites.
Hallucination or Confabulation? Neuroanatomy as metaphor in Large Language Models
A. L. Smith, F. Greaves, and T. Panch. 2023 · 2023
Later among the works it cites.
MemoryBank: Enhancing Large Language Models with Long-Term Memory
W. Zhong, L. Guo, Q. Gao, H. Ye, and Y. Wang. 2023 · 2023
Later among the works it cites.
M. S. Aissi, C. Romac, T. Carta, S. Lamprier, P. Oudeyer, O. Sigaud, L. Soulier, and N. Thome. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The continuity of context: A role for the hippocampus
A. P. Maurer and L. Nadel. 2021 · 2020
Cited alongside, same era.
Sliding-window thompson sampling for non-stationary settings
F. Trovo, S. Paladino, M. Restelli, and N. Gatti. 2020 · 2020
Cited alongside, same era.
Causal Decision Making and Causal Effect Estimation Are Not the Same…and Why It Matters
C. Fernández-Loría and F. Provost. 2022 · 2021
Cited alongside, same era.
A path towards autonomous machine intelligence
Y. LeCun. 2022 · 2022
Cited alongside, same era.
A Generalist Agent
Scott R., K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-maron, M. Giménez, Y. Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. Edwards, N. Heess, Y. Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas. 2022 · 2022
Cited alongside, same era.
Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, M. Wang, and H. Wang. 2024 · 2024
Later among the works it cites.
Understanding LLMs: A comprehensive overview from training to inference
Y. Liu, H. He, T. Han, X. Zhang, M. Liu, J. Tian, Y. Zhang, J. Wang, X. Gao, T. Zhong, Y. Pan, S. Xu, Z. Wu, Z. Liu, X. Zhang, S. Zhang, X. Hu, T. Zhang, N. Qiang, T. Liu, and B. Ge. 2025 · 2024
Later among the works it cites.
On Limitations of the Transformer Architecture. In First Conference on Language Modeling
B. Peng, S. Narayanan, and C. Papadimitriou. 2024 · 2024
Later among the works it cites.
Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
X. Wang, M. Salmani, P. Omidi, X. Ren, M. Rezagholizadeh, and A. Eshaghi. 2024 · 2024
Later among the works it cites.
Efficient Solutions For An Intriguing Failure of LLMs: Long Context Window Does Not Mean LLMs Can Analyze Long Sequences Flawlessly. In Proc.of the 31st International Conference on Computational Linguistics . Association for Computational Linguistics, 1880–1891
P. Hosseini, I. Castro, I. Ghinassi, and M. Purver. 2025 · 2025
Closest in time.
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
L. Ruis, M. Mozes, J. Bae, S. R. Kamalakara, D. Talupuru, A. Locatelli, R. Kirk, T. Rocktäschel, E. Grefenstette, and M. Bartolo. 2025 · 2025
Closest in time.