Fetching the paper…
Reading the bibliography…
Recently, large language models (LLMs) have made remarkable progress in natural language processing.
Availability versus accessibility of information in memory for words
Endel Tulving and Zena Pearlstone · 1966
Earlier work this paper cites.
Encoding specificity and retrieval processes in episodic memory
Endel Tulving and Donald M Thomson · 1973
Earlier work this paper cites.
Context-dependent memory in two natural environments: On land and underwater
Duncan R Godden and Alan D Baddeley · 1975
Earlier work this paper cites.
Infinite-ranged models of spin-glasses
Kirkpatrick Scott and Sherrington David · 1978
Earlier work this paper cites.
Remembering in and out of context
Steven M Smith · 1979
Earlier work this paper cites.
The cue-dependent nature of state-dependent retrieval
James Eric Eich · 1980
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
John J Hopfield · 1982
Earlier work this paper cites.
Extralist cuing and retrieval inhibition
Douglas L Nelson, Cathy L McEvoy, and Martha A Friedrich · 1982
Earlier work this paper cites.
Sparse distributed memory
Pentti Kanerva · 1988
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 1993
Earlier work this paper cites.
Toward optimal active learning through monte carlo estimation of error reduction
Nicholas Roy and Andrew McCallum · 2001
Earlier work this paper cites.
Prefrontal cortex and episodic memory: Integrating findings from neuropsychology and functional brain imaging
Charan Ranganath and Robert T Knight · 2002
Earlier work this paper cites.
The medial prefrontal cortex is involved in spatial memory retrieval under partial-cue conditions
Yong Sang Jo, Eun Hye Park, Il Hwan Kim, Soon Kwon Park, Hyun Kim, Hyun Taek Kim, and June-Seek Choi · 2007
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson, Hugo Zaragoza, et al · 2009
Earlier work this paper cites.
Memory retrieval in response to partial cues requires nmda receptor-dependent neurotransmission in the medial prefrontal cortex
Yong Sang Jo and June-Seek Choi · 2014
Earlier work this paper cites.
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al · 2015
Earlier work this paper cites.
Can active memory replace attention?
Łukasz Kaiser and Samy Bengio · 2016
Earlier work this paper cites.
Dense associative memory for pattern recognition
Dmitry Krotov and John J Hopfield · 2016
Earlier work this paper cites.
Ask me anything: Dynamic memory networks for natural language processing
Ankit Kumar, Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury, Ishaan Gulrajani, Victor Zhong, Romain Paulus, and Richard Socher · 2016
Earlier work this paper cites.
The colorado richly annotated full text (craft) corpus: Multi-model annotation in the biomedical domain
Kevin Bretonnel Cohen, Karin M. Verspoor, Karën Fort, Christopher S. Funk, Michael Bada, Martha Palmer, and Lawrence E. Hunter · 2017
Earlier work this paper cites.
Frustratingly short attention spans in neural language modeling
Michał Daniluk, Tim Rocktäschel, Johannes Welbl, and Sebastian Riedel · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller · 2019
Cited alongside, same era.
Prediction and memory: A predictive coding account
Helen C Barron, Ryszard Auksztulewicz, and Karl Friston · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Later among the works it cites.
Universal hopfield networks: A general framework for single-shot associative memory models
Beren Millidge, Tommaso Salvatori, Yuhang Song, Thomas Lukasiewicz, and Rafal Bogacz · 2022
Later among the works it cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Later among the works it cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu · 2022
Later among the works it cites.
Learning to retrieve prompts for in-context learning
Ohad Rubin, Jonathan Herzig, and Jonathan Berant · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, et al · 2020
Cited alongside, same era.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer · 2020
Cited alongside, same era.
Rapid encoding of musical tones discovered in whole-brain connectivity
Leonardo Bonetti, Elvira Brattico, Francesco Carlomagno, Giovanni Donati, Joana Cabral, Niels Trusbak Haumann, Gustavo Deco, Peter Vuust, and Morten L Kringelbach · 2021
Cited alongside, same era.
Attention approximates sparse distributed memory
Trenton Bricken and Cengiz Pehlevan · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Cited alongside, same era.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits · 2021
Cited alongside, same era.
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al · 2022
Later among the works it cites.
An information-theoretic approach to prompt engineering without ground truth labels
Taylor Sorensen, Joshua Robinson, Christopher Rytting, Alexander Shaw, Kyle Rogers, Alexia Delorey, Mahmoud Khalil, Nancy Fulda, and David Wingate · 2022
Later among the works it cites.
Transformer memory as a differentiable search index
Yi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, et al · 2022
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2022
Later among the works it cites.
A neural corpus indexer for document retrieval
Yujing Wang, Yingyan Hou, Haonan Wang, Ziming Miao, Shibin Wu, Qi Chen, Yuqing Xia, Chengmin Chi, Guoshuai Zhao, Zheng Liu, et al · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
An explanation of in-context learning as implicit bayesian inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2022
Later among the works it cites.
Ground-truth labels matter: A deeper look into input-label demonstrations
Kang Min Yoo, Junyeob Kim, Hyuhng Joon Kim, Hyunsoo Cho, Hwiyeol Jo, Sang-Woo Lee, Sang-goo Lee, and Taeuk Kim · 2022
Later among the works it cites.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Anonymous · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Closest in time.
On sparse modern hopfield model
Jerry Yao-Chieh Hu, Donglin Yang, Dennis Wu, Chenwei Xu, Bo-Yu Chen, and Han Liu · 2023
Closest in time.
What In-Context Learning “Learns” In-Context: Disentangling Task Recognition and Task Learning
Jane Pan · 2023
Closest in time.
On the robustness of chatgpt: An adversarial and out-of-distribution perspective
Jindong Wang, Xixu Hu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Haojun Huang, Wei Ye, Xiubo Geng, et al · 2023
Closest in time.
Compositional exemplars for in-context learning
Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu, and Lingpeng Kong · 2023
Closest in time.