Fetching the paper…
Reading the bibliography…
Scaling large language models (LLMs) leads to an emergent capacity to learn in-context from example demonstrations.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew · 1910
Earlier work this paper cites.
Syntactic structures
Noam Chomsky · 1957
Earlier work this paper cites.
A formal theory of inductive inference, part ii
Ray J. Solomonoff · 1964
Earlier work this paper cites.
Gapping and the order of constituents
John Robert Ross · 1970
Earlier work this paper cites.
The proper treatment of quantification in ordinary English
Richard Montague · 1973
Earlier work this paper cites.
Generalized phrase structure grammars, head grammars, and natural language, 1984
Carl Pollard · 1984
Earlier work this paper cites.
Natural language parsing: Tree adjoining grammars: How much context-sensitivity is required to provide reasonable structural descriptions?
Aravind K. Joshi · 1985
Earlier work this paper cites.
Evidence against the context-freeness of natural language
Stuart M. Shieber · 1985
Earlier work this paper cites.
Characterizing structural descriptions produced by various grammatical formalisms
K. Vijay-Shanker, David J. Weir, and Aravind K. Joshi · 1987
Earlier work this paper cites.
Gapping as constituent coordination
Mark Steedman · 1990
Earlier work this paper cites.
Minimum complexity density estimation
Andrew R. Barron and Thomas M. Cover · 1991
Earlier work this paper cites.
On multiple context-free grammars
Hiroyuki Seki, Takashi Matsumura, Mamoru Fujii, and Tadao Kasami · 1991
Earlier work this paper cites.
The minimalist program
Noam Chomsky · 1992
Earlier work this paper cites.
From discourse to logic - introduction to modeltheoretic semantics of natural language, formal logic and discourse representation theory
Hans Kamp and Uwe Reyle · 1993
Earlier work this paper cites.
Head-driven phrase structure grammar
Carl Pollard and Ivan A Sag · 1994
Earlier work this paper cites.
The equivalence of four extensions of context-free grammars
K. Vijay-Shanker and David J. Weir · 1994
Earlier work this paper cites.
Stochastic attribute-value grammars
Steven P. Abney · 1996
Earlier work this paper cites.
Derivational minimalism
E. Stabler · 1996
Earlier work this paper cites.
Rethinking eliminative connectionism
Gary F. Marcus · 1998
Earlier work this paper cites.
Derivational minimalism is mildly context-sensitive
Jens Michaelis · 1998
Earlier work this paper cites.
Chinese numbers, mix, scrambling, and range concatenation grammars
Pierre Boullier · 1999
Earlier work this paper cites.
Statistical properties of probabilistic context-free grammars
Zhiyi Chi · 1999
Earlier work this paper cites.
Foundations of statistical natural language processing
Christopher D. Manning and Hinrich Schütze · 1999
Earlier work this paper cites.
Range concatenation grammars
Pierre Boullier · 2000
Earlier work this paper cites.
Lexical Functional Syntax
Joan Bresnan · 2000
Earlier work this paper cites.
Interrogative Investigations: The Form, Meaning, and Use of English Interrogatives
Jonathan Ginzburg and Ivan A. Sag · 2001
Earlier work this paper cites.
Transforming linear context-free rewriting systems into minimalist grammars
Jens Michaelis · 2001
Earlier work this paper cites.
The syntactic process
Mark Steedman · 2001
Cited alongside, same era.
Universal artificial intelligence
Marcus Hutter · 2004
Cited alongside, same era.
Ltag semantics with semantic unification
Laura Kallmeyer and Maribel Romero · 2004
Cited alongside, same era.
Features moving madly: A formal perspective on feature percolation in the minimalist program
Gregory M Kobele · 2005
Cited alongside, same era.
Constructions at work: The nature of generalization in language
Adele E. Goldberg · 2006
Cited alongside, same era.
Uncertainty about the rest of the sentence
John Hale · 2006
Cited alongside, same era.
Constraints on multiple center-embedding of clauses
Visually grounded compound pcfgs
Yanpeng Zhao and Ivan Titov · 2020
Later among the works it cites.
Strong learning of some probabilistic multiple context-free grammars
Alexander Clark · 2021
Later among the works it cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell I. Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena · 2021
Later among the works it cites.
Extrapolating to unnatural language processing with gpt-3’s in-context learning: The good, the bad, and the mysterious
Frieda Rong · 2021
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fred Karlsson · 2007
Cited alongside, same era.
English Syntax: An Introduction
Jong-Bok Kim and Peter Sells · 2008
Cited alongside, same era.
An Introduction to Kolmogorov Complexity and its Applications
Ming Li and Paul Vitányi · 2008
Cited alongside, same era.
Minimum description length principle
Jorma Rissanen · 2010
Cited alongside, same era.
Sign-Based Construction Grammar
Hans Christian Boas and Ivan A. Sag · 2012
Cited alongside, same era.
The interactive stance : meaning for conversation
Jonathan Ginzburg · 2012
Cited alongside, same era.
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2022
Later among the works it cites.
Data distributional properties drive emergent in-context learning in transformers
Stephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang, Aaditya Singh, Pierre H. Richemond, Jay McClelland, and Felix Hill · 2022
Later among the works it cites.
Why can GPT learn in-context? language models secretly perform gradient descent as meta-optimizers
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei · 2022
Later among the works it cites.
What can transformers learn in-context? A case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant · 2022
Later among the works it cites.
Demystifying prompts in language models via perplexity estimation
Hila Gonen, Srini Iyer, Terra Blevins, Noah A. Smith, and Luke Zettlemoyer · 2022
Later among the works it cites.
Can language models learn from explanations in context?
Andrew K. Lampinen, Ishita Dasgupta, Stephanie C. Y. Chan, Kory W. Mathewson, Mh Tessler, Antonia Creswell, James L. McClelland, Jane Wang, and Felix Hill · 2022
Later among the works it cites.
Towards understanding grokking: An effective theory of representation learning
Ziming Liu, Ouail Kitouni, Niklas Stefan Nolte, Eric J. Michaud, Max Tegmark, and Mike Williams · 2022
Later among the works it cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Later among the works it cites.
Coloring the blank slate: Pre-training imparts a hierarchical inductive bias to sequence-to-sequence models
Aaron Mueller, Robert Frank, Tal Linzen, Luheng Wang, and Sebastian Schuster · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, T. J. Henighan, Benjamin Mann, Amanda Askell, Yushi Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, John Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom B. Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Christopher Olah · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harrison Edwards, Igor Babuschkin, and Vedant Misra · 2022
Later among the works it cites.
Impact of pretraining term frequencies on few-shot numerical reasoning
Yasaman Razeghi, Robert L. Logan IV, Matt Gardner, and Sameer Singh · 2022
Later among the works it cites.
On the effect of pretraining corpora on in-context learning by a large-scale language model
Seongjin Shin, Sang-Woo Lee, Hwijeen Ahn, Sungdong Kim, HyoungSeok Kim, Boseop Kim, Kyunghyun Cho, Gichang Lee, Woo-Myoung Park, Jung-Woo Ha, and Nako Sung · 2022
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Scharli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed Huai hsin Chi, Denny Zhou, and Jason Wei · 2022
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, Joao Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2022
Later among the works it cites.
Iteratively prompt pre-trained language models for chain of thought
Boshi Wang, Xiang Deng, and Huan Sun · 2022
Later among the works it cites.
An explanation of in-context learning as implicit bayesian inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2022
Later among the works it cites.
Unsupervised discontinuous constituency parsing with mildly context-sensitive grammars
Songlin Yang, R. Levy, and Yoon Kim · 2022
Later among the works it cites.
A toy model of universality: Reverse engineering how networks learn group operations
Bilal Chughtai, Lawrence Chan, and Neel Nanda · 2023
Closest in time.
Transformers as algorithms: Generalization and implicit model selection in in-context learning
Yingcong Li, M. Emrullah Ildiz, Dimitris S. Papailiopoulos, and Samet Oymak · 2023
Closest in time.
Xinyi Wang, Wanrong Zhu, and William Yang Wang · 2023
Closest in time.
Larger language models do in-context learning differently
Jerry W. Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, and Tengyu Ma · 2023
Closest in time.