Fetching the paper…
Reading the bibliography…
In-context learning (ICL), the remarkable ability to solve a task from only input exemplars, is often assumed to be a unique hallmark of Transformer models.
Are theories of learning necessary?
Burrhus Frederic Skinner · 1950
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Jerry A Fodor and Zenon W Pylyshyn · 1988
Earlier work this paper cites.
Rethinking eliminative connectionism
Gary F Marcus · 1998
Earlier work this paper cites.
Rule learning by seven-month-old infants
Gary F Marcus, Sugumaran Vijayan, Shoba Bandi Rao, and Peter M Vishton · 1999
Earlier work this paper cites.
Feature selection, l 1 vs. l 2 regularization, and rotational invariance
Andrew Y Ng · 2004
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Relational inductive biases, deep learning, and graph networks
Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al · 2018
Earlier work this paper cites.
Not-so-clevr: learning same–different relations strains feedforward neural networks
Junkyung Kim, Matthew Ricci, and Thomas Serre · 2018
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden Lake and Marco Baroni · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs , 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
The bitter lesson
Richard Sutton · 2019
Earlier work this paper cites.
A review of computational models of basic rule learning: The neural-symbolic debate and beyond
Raquel G Alhama and Willem Zuidema · 2019
Earlier work this paper cites.
Deep learning: the good, the bad, and the ugly
Thomas Serre · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2020
Cited alongside, same era.
pandas-dev/pandas: Pandas , February 2020
The pandas development team · 2020
Cited alongside, same era.
Sensitivity to geometric shape regularity in humans and baboons: A putative signature of human singularity
Mathias Sablé-Meyer, Joël Fagot, Serge Caparos, Timo van Kerkoerle, Marie Amalric, and Stanislas Dehaene · 2021
Cited alongside, same era.
Mlp-mixer: An all-mlp architecture for vision
Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al · 2021
Cited alongside, same era.
Pay attention to mlps
Hanxiao Liu, Zihang Dai, David So, and Quoc V Le · 2021
Cited alongside, same era.
Declan Campbell, Sreejan Kumar, Tyler Giallanza, Jonathan D Cohen, and Thomas L Griffiths · 2023
Later among the works it cites.
Transformers as algorithms: Generalization and stability in in-context learning
Yingcong Li, Muhammed Emrullah Ildiz, Dimitris Papailiopoulos, and Samet Oymak · 2023
Later among the works it cites.
Relational reasoning and generalization using nonsymbolic neural networks
Atticus Geiger, Alexandra Carstensen, Michael C Frank, and Christopher Potts · 2023
Later among the works it cites.
The relational bottleneck as an inductive bias for efficient abstraction
Taylor W Webb, Steven M Frankland, Awni Altabaa, Kamesh Krishnamurthy, Declan Campbell, Jacob Russin, Randall O’Reilly, John Lafferty, and Jonathan D Cohen · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An explanation of in-context learning as implicit bayesian inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2021
Cited alongside, same era.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Cited alongside, same era.
seaborn: statistical data visualization
Michael L. Waskom · 2021
Cited alongside, same era.
A survey on in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant · 2022
Cited alongside, same era.
Data distributional properties drive emergent in-context learning in transformers
Stephanie Chan, Adam Santoro, Andrew Lampinen, Jane Wang, Aaditya Singh, Pierre Richemond, James McClelland, and Felix Hill · 2022
Cited alongside, same era.
pnlp-mixer: An efficient all-mlp architecture for language
Francesco Fusco, Damian Pascual, Peter Staar, and Diego Antognini · 2022
Cited alongside, same era.
Enric Boix-Adsera, Omid Saremi, Emmanuel Abbe, Samy Bengio, Etai Littwin, and Joshua Susskind · 2023
Later among the works it cites.
Some intriguing aspects about lipschitz continuity of neural networks
Grigory Khromov and Sidak Pal Singh · 2023
Later among the works it cites.
Flax: A neural network library and ecosystem for JAX , 2023
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee · 2023
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2024
Closest in time.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task
Gautam Reddy · 2024
Closest in time.
Asymptotic theory of in-context learning by linear attention
Yue M. Lu, Mary I. Letey, Jacob A. Zavatone-Veth, Anindita Maiti, and Cengiz Pehlevan · 2024
Closest in time.
Pretraining task diversity and the emergence of non-bayesian in-context learning for regression
Allan Raventós, Mansheej Paul, Feng Chen, and Surya Ganguli · 2024
Closest in time.
Scaling mlps: A tale of inductive bias
Gregor Bachmann, Sotiris Anagnostidis, and Thomas Hofmann · 2024
Closest in time.
How many pretraining tasks are needed for in-context learning of linear regression?
Jingfeng Wu, Difan Zou, Zixiang Chen, Vladimir Braverman, Quanquan Gu, and Peter L Bartlett · 2024
Closest in time.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2024
Closest in time.
In-context learning through the bayesian prism
Kabir Ahuja, Madhur Panwar, and Navin Goyal · 2024
Closest in time.
Exploring the relationship between model architecture and in-context learning ability
Ivan Lee, Nan Jiang, and Taylor Berg-Kirkpatrick · 2024
Closest in time.
Can mamba learn how to learn? a comparative study on in-context learning tasks
Jongho Park, Jaeseung Park, Zheyang Xiong, Nayoung Lee, Jaewoong Cho, Samet Oymak, Kangwook Lee, and Dimitris Papailiopoulos · 2024
Closest in time.