Fetching the paper…
Reading the bibliography…
In-context learning (ICL) is one of the most powerful and most unexpected capabilities to emerge in recent transformer-based large language models (LLMs).
“Language models are few-shot learners”
Tom Brown et al · 1901
Earlier work this paper cites.
“Continuous speech recognition by statistical methods”
Frederick Jelinek · 1976
Earlier work this paper cites.
“Maximum likelihood from incomplete data via the EM algorithm”
Arthur Dempster, Nan Laird and Donald Rubin · 1977
Earlier work this paper cites.
“Learning and reasoning by analogy”
Patrick Winston · 1980
Earlier work this paper cites.
“Probabilistic reasoning in intelligent systems: networks of plausible inference”
Judea Pearl · 1988
Earlier work this paper cites.
“An algorithm for drawing general undirected graphs”
Tomihisa Kamada and Satoru Kawai · 1989
Earlier work this paper cites.
“Reinforcement learning with perceptual aliasing: The perceptual distinctions approach”
Lonnie Chrisman · 1992
Earlier work this paper cites.
“Meta-neural networks that learn by learning”
Devang Naik and Richard Mammone · 1992
Earlier work this paper cites.
“Factorial hidden Markov models”
Zoubin Ghahramani and Michael Jordan · 1995
Earlier work this paper cites.
“Learning to learn using gradient descent”
Sepp Hochreiter, A Younger and Peter Conwell · 2001
Earlier work this paper cites.
“One-shot learning of object categories”
Li Fei-Fei, Robert Fergus and Pietro Perona · 2006
Earlier work this paper cites.
“Exact Bayesian structure learning from uncertain interventions”
Daniel Eaton and Kevin Murphy · 2007
Earlier work this paper cites.
“Causality”
Judea Pearl · 2009
Earlier work this paper cites.
“Adam: A method for stochastic optimization”
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
“Learning a generative probabilistic grammar of experience: A process-level model of language acquisition”
Oren Kolodny, Arnon Lotem and Shimon Edelman · 2015
Earlier work this paper cites.
“Human-level concept learning through probabilistic program induction”
Brenden Lake, Ruslan Salakhutdinov and Joshua Tenenbaum · 2015
Earlier work this paper cites.
“An overview of gradient descent optimization algorithms”
Sebastian Ruder · 2016
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Learning overcomplete hmms”
Vatsal Sharan, Sham Kakade, Percy Liang and Gregory Valiant · 2017
Earlier work this paper cites.
“Elements of causal inference: foundations and learning algorithms”
Jonas Peters, Dominik Janzing and Bernhard Schölkopf · 2017
Cited alongside, same era.
“Remember dax? Relations between children’s cross-situational word learning, memory, and language abilities”
Haley Vlach and Catherine DeBrock · 2017
Cited alongside, same era.
“Behavioral time scale synaptic plasticity underlies CA1 place fields”
Katie Bittner et al · 2017
Cited alongside, same era.
“Optimization as a model for few-shot learning”
Sachin Ravi and Hugo Larochelle · 2017
Cited alongside, same era.
“PreCo: A large-scale dataset in preschool vocabulary for coreference resolution”
Hong Chen et al · 2018
Cited alongside, same era.
“Optimization methods for large-scale machine learning”
“Why Can GPT Learn In-Context? Language Models Secretly Perform Gradient Descent as Meta Optimizers”
Damai Dai et al · 2022
Later among the works it cites.
“In-context learning and induction heads”
Catherine Olsson et al · 2022
Later among the works it cites.
“Abstraction for Deep Reinforcement Learning”
Murray Shanahan and Melanie Mitchell · 2022
Later among the works it cites.
“Locating and Editing Factual Associations in GPT”
Kevin Meng, David Bau, Alex Andonian and Yonatan Belinkov · 2022
Later among the works it cites.
“Transformers learn shortcuts to automata”
Bingbin Liu et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Léon Bottou, Frank Curtis and Jorge Nocedal · 2018
Cited alongside, same era.
“Learning higher-order sequential structure with cloned HMMs”
Antoine Dedieu et al · 2019
Cited alongside, same era.
“Generalizing from a few examples: A survey on few-shot learning”
Yaqing Wang, Quanming Yao, James Kwok and Lionel Ni · 2020
Cited alongside, same era.
“An explanation of in-context learning as implicit bayesian inference”
Sang Xie, Aditi Raghunathan, Percy Liang and Tengyu Ma · 2021
Cited alongside, same era.
“Clone-structured graph representations enable flexible learning and vicarious evaluation of cognitive maps”
Dileep George et al · 2021
Cited alongside, same era.
“Multitask prompted training enables zero-shot task generalization”
Victor Sanh et al · 2021
Cited alongside, same era.
“Finetuned language models are zero-shot learners”
Jason Wei et al · 2021
Cited alongside, same era.
Jason Wei et al · 2022
Later among the works it cites.
“Least-to-most prompting enables complex reasoning in large language models”
Denny Zhou et al · 2022
Later among the works it cites.
“Data distributional properties drive emergent in-context learning in transformers”
Stephanie Chan et al · 2022
Later among the works it cites.
“What can transformers learn in-context? a case study of simple function classes”
Shivam Garg, Dimitris Tsipras, Percy Liang and Gregory Valiant · 2022
Later among the works it cites.
“What learning algorithm is in-context learning? investigations with linear models”
Ekin Akyürek et al · 2022
Later among the works it cites.
“Systematic Generalization and Emergent Structures in Transformers Trained on Structured Tasks”
Yuxuan Li and James McClelland · 2022
Later among the works it cites.
“On the effect of pretraining corpora on in-context learning by a large-scale language model”
Seongjin Shin et al · 2022
Later among the works it cites.
“Emergent abilities of large language models”
Jason Wei et al · 2022
Later among the works it cites.
“Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?”
Sewon Min et al · 2022
Later among the works it cites.
“Can language models learn from explanations in context?”
Andrew Lampinen et al · 2022
Later among the works it cites.
“Progress measures for grokking via mechanistic interpretability”
Neel Nanda et al · 2023
Closest in time.
“Graph schemas as abstractions for transfer learning, inference, and planning”
J Guntupalli et al · 2023
Closest in time.
“Are Emergent Abilities of Large Language Models a Mirage?”
Rylan Schaeffer, Brando Miranda and Sanmi Koyejo · 2023
Closest in time.