Fetching the paper…
Reading the bibliography…
Transformers have the capacity to act as supervised learning algorithms: by properly encoding a set of labeled training ("in-context") examples and an unlabeled test example into an input sequence of vectors of the same dimension, the forward pass of the transformer can produce predictions for that unlabeled test example.
“Language models are few-shot learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry and Amanda Askell · 1901
Earlier work this paper cites.
“A note on the Hanson-Wright inequality for random vectors with dependencies”
Radoslaw Adamczak · 2015
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Understanding deep learning requires rethinking generalization”
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht and Oriol Vinyals · 2017
Earlier work this paper cites.
“The implicit bias of gradient descent on separable data”
Daniel Soudry, Elad Hoffer, Mor Nacson, Suriya Gunasekar and Nathan Srebro · 2018
Earlier work this paper cites.
“High-dimensional probability: An introduction with applications in data science”
Roman Vershynin · 2018
Earlier work this paper cites.
“Does data interpolation contradict statistical optimality?”
Mikhail Belkin, Alexander Rakhlin and Alexandre. Tsybakov · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei and Ilya Sutskever · 2019
Earlier work this paper cites.
“Benign Overfitting in Linear Regression”
Peter. Bartlett, Philip. Long, Gabor Lugosi and Alexander Tsigler · 2020
Earlier work this paper cites.
“Directional convergence and alignment in deep learning”
Ziwei Ji and Matus Telgarsky · 2020
Earlier work this paper cites.
“Gradient descent maximizes the margin of homogeneous neural networks”
Kaifeng Lyu and Jian Li · 2020
Earlier work this paper cites.
“Deep learning: a statistical viewpoint”
Peter. Bartlett, Andrea Montanari and Alexander Rakhlin · 2021
Earlier work this paper cites.
“Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation”
Mikhail Belkin · 2021
Cited alongside, same era.
“Finite-sample analysis of interpolating linear classifiers in the overparameterized regime”
Niladri. Chatterji and Philip. Long · 2021
Cited alongside, same era.
“What learning algorithm is in-context learning? investigations with linear models”
Ekin Akyurek, Dale Schuurmans, Jacob Andreas, Tengyu Ma and Denny Zhou · 2022
Cited alongside, same era.
“Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data”
Spencer Frei, Niladri. Chatterji and Peter. Bartlett · 2022
Cited alongside, same era.
“What can transformers learn in-context? a case study of simple function classes”
Shivam Garg, Dimitris Tsipras, Percy Liang and Gregory Valiant · 2022
“The Double-Edged Sword of Implicit Bias: Generalization vs. Robustness in ReLU Networks”
Spencer Frei, Gal Vardi, Peter. Bartlett and Nathan Srebro · 2023
Later among the works it cites.
“Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data”
Spencer Frei, Gal Vardi, Peter. Bartlett, Nathan Srebro and Wei Hu · 2023
Later among the works it cites.
“Transformers as support vector machines”
Davoud Tarzanagh, Yingcong Li, Christos Thrampoulidis and Samet Oymak · 2023
Later among the works it cites.
“Transformers as statisticians: Provable in-context learning with in-context algorithm selection”
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong and Song Mei · 2024
Closest in time.
“The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains”
Benjamin. Edelman, Ezra Edelman, Surbhi Goel, Eran Malach and Nikolaos Tsilivis · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Transformers learn in-context by gradient descent”
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov and Max Vladymyrov · 2022
Cited alongside, same era.
Itay Safran, Gal Vardi and Jason Lee · 2022
Cited alongside, same era.
“On the Implicit Bias in Deep-Learning Algorithms”
Gal Vardi · 2022
Cited alongside, same era.
“Transformers learn to implement preconditioned gradient descent for in-context learning”
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand and Suvrit Sra · 2023
Cited alongside, same era.
“Max-margin token selection in attention mechanism”
Davoud Ataee, Yingcong Li, Xuechen Zhang and Samet Oymak · 2023
Cited alongside, same era.
“Benign Overfitting in Linear Classifiers and Leaky ReLU Networks from KKT Conditions for Margin Maximization”
Spencer Frei, Gal Vardi, Peter. Bartlett and Nathan Srebro · 2023
Cited alongside, same era.
Closest in time.
“Transformers are Minimax Optimal Nonparametric In-Context Learners”
Juno Kim, Tai Nakamaki and Taiji Suzuki · 2024
Closest in time.
“One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention”
Arvind Mahankali, Tatsunori. Hashimoto and Tengyu Ma · 2024
Closest in time.
“Implicit Bias of Next-Token Prediction”
Christos Thrampoulidis · 2024
Closest in time.
“Implicit bias and fast convergence rates for self-attention”
Bhavya Vasudeva, Puneesh Deora and Christos Thrampoulidis · 2024
Closest in time.
“How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?”
Jingfeng Wu, Difan Zou, Zixiang Chen, Vladimir Braverman, Quanquan Gu and Peter. Bartlett · 2024
Closest in time.
“Trained Transformers Learn Linear Models In-Context”
Ruiqi Zhang, Spencer Frei and Peter. Bartlett · 2024
Closest in time.