Fetching the paper…
Reading the bibliography…
We analyze how well pre-trained large language models (e.g., Llama2, GPT-4, Claude 3, etc) can do linear and non-linear regression when given in-context examples, without any additional training or gradient updates.
Multiple comparisons among means
Olive Jean Dunn · 1961
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George V. Cybenko · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
K. Hornik, M. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
The strength of weak learnability
Robert E. Schapire · 1989
Earlier work this paper cites.
Liver Disorders
UCI · 1990
Earlier work this paper cites.
Multivariate Adaptive Regression Splines
Jerome H. Friedman · 1991
Earlier work this paper cites.
Controlling the false discovery rate: A practical and powerful approach to multiple testing
Yoav Benjamini and Yosef Hochberg · 1995
Earlier work this paper cites.
Bagging predictors
Leo Breiman · 1996
Earlier work this paper cites.
Random forests
L. Breiman · 2001
Earlier work this paper cites.
Greedy function approximation: A gradient boosting machine
Jerome H. Friedman · 2001
Earlier work this paper cites.
Least angle regression
Bradley Efron, Trevor Hastie, Iain Johnstone, and Robert Tibshirani · 2004
Earlier work this paper cites.
The elements of statistical learning: Data mining, inference, and prediction
David Ruppert · 2004
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
CSM (Conventional and Social Media Movies) Dataset 2014 and 2015
Mehreen Ahmed · 2017
Earlier work this paper cites.
The expressive power of neural networks: A view from the width
Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang · 2017
Earlier work this paper cites.
Neural arithmetic logic units
Andrew Trask, Felix Hill, Scott E Reed, Jack Rae, Chris Dyer, and Phil Blunsom · 2018
Earlier work this paper cites.
Real Estate Valuation
I-Cheng Yeh · 2018
Earlier work this paper cites.
Online meta-learning
Chelsea Finn, Aravind Rajeswaran, Sham Kakade, and Sergey Levine · 2019
Earlier work this paper cites.
A modern introduction to online learning
Francesco Orabona · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Neural power units
Niklas Heim, Tomas Pevny, and Vasek Smidl · 2020
Earlier work this paper cites.
Data distributional properties drive emergent in-context learning in transformers, 2022
Stephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang, Aaditya Singh, Pierre H. Richemond, Jay McClelland, and Felix Hill · 2022
Earlier work this paper cites.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant · 2022
Earlier work this paper cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Cited alongside, same era.
A primer for neural arithmetic logic modules
Bhumika Mistry, Katayoun Farrahi, and Jonathon Hare · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, and Pamela et al. Mishkin · 2022
Cited alongside, same era.
Impact of pretraining term frequencies on few-shot numerical reasoning
Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh · 2022
Cited alongside, same era.
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, E. Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2022
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
What in-context learning “learns” in-context: Disentangling task recognition and task learning
Jane Pan, Tianyu Gao, Howard Chen, and Danqi Chen · 2023
Later among the works it cites.
RWKV: Reinventing RNNs for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Leon Derczynski, Xingjian Du, Matteo Grella, Kranthi Gv, Xuzheng He, Haowen Hou, Przemyslaw Kazienko, Jan Kocon, Jiaming Kong, Bartłomiej Koptyra, Hayden Lau, Jiaju Lin, Krishna Sri Ipsit Mantri, Ferdinand Mom, Atsushi Saito, Guangyu Song, Xiangru Tang, Johan Wind, Stanisław Woźniak, Zhenyuan Zhang, Qinghua Zhou, Jian Zhu, and Rui-Jie Zhu · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, and Xiaoqing Tan et al · 2023
Later among the works it cites.
NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark
Oscar Sainz, Jon Campos, Iker García-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2022
Cited alongside, same era.
Gpt-4 technical report
OpenAI Josh Achiam, Steven Adler, Sandhini Agarwal, and Lama Ahmad et al · 2023
Cited alongside, same era.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra · 2023
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2023
Cited alongside, same era.
The falcon series of open language models
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra-Aimée Cojocaru, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo · 2023
Cited alongside, same era.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2023
Cited alongside, same era.
Transformers implement functional gradient descent to learn non-linear functions in context
Xiang Cheng, Yuxin Chen, and Suvrit Sra · 2023
Cited alongside, same era.
Lingfeng Shen, Aayush Mishra, and Daniel Khashabi · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2023
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, and Jean-Baptiste Alayrac et al · 2023
Later among the works it cites.
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer
Yuandong Tian, Yiping Wang, Beidi Chen, and Simon Shaolei Du · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, and Yasmine Babaei et al · 2023
Later among the works it cites.
Uncovering mesa-optimization algorithms in transformers
Johannes von Oswald, Eyvind Niklasson, Maximilian Schlegel, Seijin Kobayashi, Nicolas Zucchet, Nino Scherrer, Nolan Miller, Mark Sandler, Blaise Agüera y Arcas, Max Vladymyrov, Razvan Pascanu, and João Sacramento · 2023
Later among the works it cites.
Larger language models do in-context learning differently
Jerry W. Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, and Tengyu Ma · 2023
Later among the works it cites.
The learnability of in-context learning
Noam Wies, Yoav Levine, and Amnon Shashua · 2023
Later among the works it cites.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter Bartlett · 2023
Later among the works it cites.
Yi: Open foundation models by 01.ai, 2024
01. AI, Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, and Guanwei Zhang et al · 2024
Closest in time.
Understanding in-context learning in transformers and LLMs by learning to learn discrete functions
Satwik Bhattamishra, Arkil Patel, Phil Blunsom, and Varun Kanade · 2024
Closest in time.
Time travel in LLMs: Tracing data contamination in large language models
Shahriar Golchin and Mihai Surdeanu · 2024
Closest in time.
How do transformers learn in-context beyond simple functions? a case study on learning with representations
Tianyu Guo, Wei Hu, Song Mei, Huan Wang, Caiming Xiong, Silvio Savarese, and Yu Bai · 2024
Closest in time.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, and Chris Bamford et al · 2024
Closest in time.
In-context learning learns label relationships but is not conventional learning
Jannik Kossen, Yarin Gal, and Tom Rainforth · 2024
Closest in time.
Hongkang Li, Meng Wang, Songtao Lu, Xiaodong Cui, and Pin-Yu Chen · 2024
Closest in time.
One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention
Arvind V. Mahankali, Tatsunori Hashimoto, and Tengyu Ma · 2024
Closest in time.
Linear transformers are versatile in-context learners
Max Vladymyrov, Johannes von Oswald, Mark Sandler, and Rong Ge · 2024
Closest in time.
Benefits of transformer: In-context learning in linear regression tasks with unstructured data
Yue Xing, Xiaofeng Lin, Namjoon Suh, Qifan Song, and Guang Cheng · 2024
Closest in time.