Fetching the paper…
Reading the bibliography…
Language models have the ability to perform in-context learning (ICL), allowing them to flexibly adapt their behavior based on context.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J. Cohen · 1989
Earlier work this paper cites.
Connectionist models of recognition memory: constraints imposed by learning and forgetting functions
Roger Ratcliff · 1990
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
When bert forgets how to POS: amnesic probing of linguistic properties and MLM predictions
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg · 2006
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman · 2011
Earlier work this paper cites.
CoNLL-2012 shared task: Modeling multilingual unrestricted coreference in OntoNotes
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Olga Uryupina, and Yuchen Zhang · 2012
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Measuring catastrophic forgetting in neural networks
Ronald Kemker, Angelina Abitino, Marc McClure, and Christopher Kanan · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2017
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 2017
Earlier work this paper cites.
Universal Dependency annotation for multilingual parsing
Ryan McDonald, Joakim Nivre, Yvonne Quirmbach-Brundage, Yoav Goldberg, Dipanjan Das, Kuzman Ganchev, Keith Hall, Slav Petrov, Hao Zhang, Oscar Täckström, Claudia Bedini, Núria Bertomeu Castelló, and Jungmee Lee · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning · 2019
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Cited alongside, same era.
Well-read students learn better: On the importance of pre-training compact models, 2019
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2020
Cited alongside, same era.
The impact of reinitialization on generalization in convolutional neural networks, 2021
Ibrahim Alabdulmohsin, Hartmut Maennel, and Daniel Keysers · 2021
Cited alongside, same era.
Neural collapse in the intermediate hidden layers of classification neural networks, 2023
Liam Parker, Emre Onal, Anton Stengel, and Jake Intrater · 2023
Later among the works it cites.
Learn, unlearn and relearn: An online learning paradigm for deep neural networks, 2023
Vijaya Raghavan T. Ramkumar, Elahe Arani, and Bahram Zonooz · 2023
Later among the works it cites.
Feature learning in deep classifiers through intermediate neural collapse
Akshay Rangamani, Marius Lindegaard, Tomer Galanti, and Tomaso A Poggio · 2023
Later among the works it cites.
Generalization to new sequential decision making tasks with in-context learning, 2023
Sharath Chandra Raparthy, Eric Hambro, Robert Kirk, Mikael Henaff, and Roberta Raileanu · 2023
Later among the works it cites.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task, 2023
Gautam Reddy · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Hewitt, Kawin Ethayarajh, Percy Liang, and Christopher Manning · 2021
Cited alongside, same era.
The multiberts: BERT reproductions for robustness analysis
Thibault Sellam, Steve Yadlowsky, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das, Ian Tenney, and Ellie Pavlick · 2021
Cited alongside, same era.
Knowledge evolution in neural networks, 2021
Ahmed Taha, Abhinav Shrivastava, and Larry Davis · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2022
Cited alongside, same era.
Fortuitous forgetting in connectionist networks, 2022
Hattie Zhou, Ankit Vani, Hugo Larochelle, and Aaron Courville · 2022
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Cited alongside, same era.
A survey on in-context learning, 2023
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, Lei Li, and Zhifang Sui · 2023
Cited alongside, same era.
Solidgoldmagikarp (plus, prompt generation)
Jessica Rumbelow and Matthew Watkins · 2023
Later among the works it cites.
The transient nature of emergent in-context learning in transformers
Aaditya K Singh, Stephanie C.Y. Chan, Ted Moskovitz, Erin Grant, Andrew M Saxe, and Felix Hill · 2023
Later among the works it cites.
In-context language learning: Architectures and algorithms, 2024
Ekin Akyürek, Bailin Wang, Yoon Kim, and Jacob Andreas · 2024
Closest in time.
Neural networks learn statistics of increasing complexity, 2024
Nora Belrose, Quintin Pope, Lucia Quirke, Alex Mallen, and Xiaoli Fern · 2024
Closest in time.
Improving language plasticity via pretraining with active forgetting, 2024
Yihong Chen, Kelly Marchisio, Roberta Raileanu, David Ifeoluwa Adelani, Pontus Stenetorp, Sebastian Riedel, and Mikel Artetxe · 2024
Closest in time.
How does representation impact in-context learning: An exploration on a synthetic task, 2024
Jingwen Fu, Tao Yang, Yuwang Wang, Yan Lu, and Nanning Zheng · 2024
Closest in time.
A mechanistic interpretation of syllogistic reasoning in auto-regressive language models
Geonhee Kim, Marco Valentino, and André Freitas · 2024
Closest in time.
Language models, like humans, show content effects on reasoning tasks
Andrew K Lampinen, Ishita Dasgupta, Stephanie C Y Chan, Hannah R Sheahan, Antonia Creswell, Dharshan Kumaran, James L McClelland, and Felix Hill · 2024
Closest in time.
Fishing for magikarp: Automatically detecting under-trained tokens in large language models, 2024
Sander Land and Max Bartolo · 2024
Closest in time.