Fetching the paper…
Reading the bibliography…
Despite their impressive performance on complex tasks, current language models (LMs) typically operate in a vacuum: Each input query is processed separately, without retaining insights from previous attempts.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M. and Cohen, N. J · 1989
Earlier work this paper cites.
Lifelong robot learning
Thrun, S. and Mitchell, T. M · 1995
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I · 1998
Earlier work this paper cites.
Handling concept drifts in incremental learning with support vector machines
Syed, N. A., Liu, H., and Sung, K. K · 1999
Earlier work this paper cites.
Large scale online learning
Bottou, L. and Cun, Y · 2003
Earlier work this paper cites.
On-line learning for very large data sets
Bottou, L. and Le Cun, Y · 2005
Earlier work this paper cites.
The critical importance of retrieval for learning
Karpicke, J. D. and Roediger III, H. L · 2008
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Cernockỳ, J., and Khudanpur, S · 2010
Earlier work this paper cites.
Retrieval practice produces more learning than elaborative studying with concept mapping
Karpicke, J. D. and Blunt, J. R · 2011
Earlier work this paper cites.
The critical role of retrieval practice in long-term retention
Roediger, H. L. and Butler, A. C · 2011
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Graves, A · 2013
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Weston, J., Chopra, S., and Bordes, A · 2014
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
Joulin, A. and Mikolov, T · 2015
Earlier work this paper cites.
Gradient episodic memory for continual learning
Lopez-Paz, D. and Ranzato, M · 2017
Earlier work this paper cites.
Dynamic evaluation of transformer language models
Krause, B., Kahembwe, E., Murray, I., and Renals, S · 2019
Earlier work this paper cites.
Metalearned neural memory
Munkhdalai, T., Sordoni, A., Wang, T., and Trischler, A · 2019
Earlier work this paper cites.
Memory-augmented recurrent neural networks can learn generalized dyck languages
Suzgun, M., Gehrmann, S., Belinkov, Y., and Shieber, S. M · 2019
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., and Smith, N. A · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Karpukhin, V., Oguz, B., Min, S., Lewis, P. S., Wu, L., Edunov, S., Chen, D., and Yih, W.-t · 2020
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Test-time training with self-supervision for generalization under distribution shifts
Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A., and Hardt, M · 2020
Earlier work this paper cites.
Tent: Fully test-time adaptation by entropy minimization
Wang, D., Shelhamer, E., Liu, S., Olshausen, B., and Darrell, T · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Cited alongside, same era.
Ttt++: When does self-supervised test-time training fail or thrive?
Liu, Y., Kothari, P., Van Delft, B., Bellot-Gurlet, B., Mordan, T., and Alahi, A · 2021
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G. B., Lespiau, J.-B., Damoc, B., Clark, A., et al · 2022
Cited alongside, same era.
Parameter-free online test-time adaptation
Boudiaf, M., Mueller, R., Ben Ayed, I., and Bertinetto, L · 2022
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Later among the works it cites.
HuggingGPT: Solving AI tasks with chatGPT and its friends in hugging face
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y · 2023
Later among the works it cites.
Language models are multilingual chain-of-thought reasoners
Shi, F., Suzgun, M., Freitag, M., Wang, X., Srivats, S., Vosoughi, S., Chung, H. W., Tay, Y., Ruder, S., Zhou, D., Das, D., and Wei, J · 2023
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2023
Later among the works it cites.
Vipergpt: Visual inference via python execution for reasoning
Surís, D., Menon, S., and Vondrick, C · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Recurrent memory transformer
Bulatov, A., Kuratov, Y., and Burtsev, M · 2022
Cited alongside, same era.
Learn to remember: Transformer with recurrent memory for document-level machine translation
Feng, Y., Li, F., Song, Z., Zheng, B., and Koehn, P · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Cited alongside, same era.
Memory-assisted prompt editing to improve gpt-3 after deployment
Madaan, A., Tandon, N., Clark, P., and Yang, Y · 2022
Cited alongside, same era.
Efficient test-time model adaptation without forgetting
Niu, S., Wu, J., Zhang, Y., Chen, Y., Zheng, S., Zhao, P., and Tan, M · 2022
Cited alongside, same era.
Natural language to code translation with execution
Shi, F., Fried, D., Ghazvininejad, M., Zettlemoyer, L., and Wang, S. I · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Suzgun, M., Melas-Kyriazi, L., and Jurafsky, D · 2023
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q., Chi, E., Zhou, D., et al · 2023
Later among the works it cites.
Freshllms: Refreshing large language models with search engine augmentation
Vu, T., Iyyer, M., Wang, X., Constant, N., Wei, J., Wei, J., Tar, C., Sung, Y.-H., Zhou, D., Le, Q., et al · 2023
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D · 2023
Later among the works it cites.
Tree of Thoughts: Deliberate problem solving with large language models, 2023
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Later among the works it cites.
Graph of thoughts: Solving elaborate problems with large language models
Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Podstawski, M., Gianinazzi, L., Gajda, J., Lehmann, T., Niewiadomski, H., Nyczyk, P., et al · 2024
Later among the works it cites.
Self-play fine-tuning converts weak language models to strong language models
Chen, Z., Deng, Y., Yuan, H., Ji, K., and Gu, Q · 2024
Later among the works it cites.
Thought-retriever: Don’t just retrieve raw data, retrieve thoughts, 2024
Feng, T., Han, P., Lin, G., Liu, G., and You, J · 2024
Later among the works it cites.
Camelot: Towards large language models with training-free consolidated associative memory
He, Z., Karlinsky, L., Kim, D., McAuley, J., Krotov, D., and Feris, R · 2024
Later among the works it cites.
Revisiting dynamic evaluation: Online adaptation for large language models
Rannen-Triki, A., Bornschein, J., Pascanu, R., Hutter, M., György, A., Galashov, A., Teh, Y. W., and Titsias, M. K · 2024
Later among the works it cites.
GPQA: A graduate-level google-proof q&a benchmark
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2024
Later among the works it cites.
REPLUG: Retrieval-augmented black-box language models
Shi, W., Min, S., Yasunaga, M., Seo, M., James, R., Lewis, M., Zettlemoyer, L., and Yih, W.-t · 2024
Later among the works it cites.
Learning to (learn at test time): Rnns with expressive hidden states
Sun, Y., Li, X., Dalal, K., Xu, J., Vikram, A., Zhang, G., Dubois, Y., Chen, X., Wang, X., Koyejo, S., et al · 2024
Later among the works it cites.
Meta-prompting: Enhancing language models with task-agnostic scaffolding
Suzgun, M. and Kalai, A. T · 2024
Later among the works it cites.
string2string: A modern python library for string-to-string algorithms
Suzgun, M., Shieber, S. M., and Jurafsky, D · 2024
Later among the works it cites.
LLM-based medical assistant personalization with short- and long-term memory coordination
Zhang, K., Kang, Y., Zhao, F., and Liu, X · 2024
Later among the works it cites.
Chain-of-thought reasoning in the wild is not always faithful
Arcuschin, I., Janiak, J., Krzyzanowski, R., Rajamanoharan, S., Nanda, N., and Conmy, A · 2025
Closest in time.
Buffer of thoughts: Thought-augmented reasoning with large language models
Yang, L., Yu, Z., Zhang, T., Cao, S., Xu, M., Zhang, W., Gonzalez, J. E., and Cui, B · 2025
Closest in time.
Optimizing generative ai by backpropagating language model feedback
Yuksekgonul, M., Bianchi, F., Boen, J., Liu, S., Lu, P., Huang, Z., Guestrin, C., and Zou, J · 2025
Closest in time.