Fetching the paper…
Reading the bibliography…
Transformers have demonstrated exceptional in-context learning capabilities, yet the theoretical understanding of the underlying mechanisms remains limited.
A mathematical theory of communication
Claude Elwood Shannon · 1948
Earlier work this paper cites.
Methods of modern mathematical physics: Functional analysis , volume 1
Michael Reed and Barry Simon · 1980
Earlier work this paper cites.
Neural net approximation
Andrew R Barron · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R. Barron · 1993
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R Barron · 1994
Earlier work this paper cites.
Generalized gaussian quadratures and singular value decompositions of integral operators
Norman Yarvin and Vladimir Rokhlin · 1998
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Earlier work this paper cites.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2019
Earlier work this paper cites.
A priori estimates of the population risk for two-layer neural networks
Weinan E, Chao Ma, and Lei Wu · 2019
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J Reddi, and Sanjiv Kumar · 2019
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Earlier work this paper cites.
Chao Ma, Stephan Wojtowytsch, Lei Wu, and Weinan E · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Nonparametric regression using deep neural networks with ReLU activation function
Johannes Schmidt-Hieber et al · 2020
Earlier work this paper cites.
Approximation rates for neural networks with general activation functions
Jonathan W Siegel and Jinchao Xu · 2020
Earlier work this paper cites.
The barron space and the flow-induced function spaces for neural network models
Weinan E, Chao Ma, and Lei Wu · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Earlier work this paper cites.
Attention is turing complete
Jorge Pérez, Pablo Barceló, and Javier Marinkovic · 2021
Cited alongside, same era.
Thinking like transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2022
Cited alongside, same era.
Kerple: Kernelized relative positional embedding for length extrapolation
Ta-Chung Chi, Ting-Han Fan, Peter J Ramadge, and Alexander Rudnicky · 2022
Cited alongside, same era.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer
Yuandong Tian, Yiping Wang, Beidi Chen, and Simon S Du · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Later among the works it cites.
Learning hierarchical polynomials with three-layer neural networks
Zihao Wang, Eshaan Nichani, and Jason D Lee · 2023
Later among the works it cites.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L Bartlett · 2023
Later among the works it cites.
Transformers learn through gradual rank increase
Emmanuel Abbe, Samy Bengio, Enric Boix-Adsera, Etai Littwin, and Joshua Susskind · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant · 2022
Cited alongside, same era.
Saturated transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A Smith · 2022
Cited alongside, same era.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A Smith, and Mike Lewis · 2022
Cited alongside, same era.
Optimization-based separations for neural networks
Itay Safran and Jason Lee · 2022
Cited alongside, same era.
Statistically meaningful approximation: a case study on approximating turing machines with transformers
Colin Wei, Yining Chen, and Tengyu Ma · 2022
Cited alongside, same era.
Max-margin token selection in attention mechanism
Davoud Ataee Tarzanagh, Yingcong Li, Xuechen Zhang, and Samet Oymak · 2023
Cited alongside, same era.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra · 2024
Closest in time.
Birth of a transformer: A memory viewpoint
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, and Leon Bottou · 2024
Closest in time.
What can transformer learn with varying depth? case studies on sequence learning tasks
Xingwu Chen and Difan Zou · 2024
Closest in time.
Induction heads as an essential mechanism for pattern matching in in-context learning
Joy Crosbie and Ekaterina Shutova · 2024
Closest in time.
The evolution of statistical induction heads: In-context learning markov chains
Benjamin L Edelman, Ezra Edelman, Surbhi Goel, Eran Malach, and Nikolaos Tsilivis · 2024
Closest in time.
Repeat after me: Transformers are better than state space models at copying
Samy Jelassi, David Brandfonbrener, Sham M Kakade, and Eran Malach · 2024
Closest in time.
Hongkang Li, Meng Wang, Songtao Lu, Xiaodong Cui, and Pin-Yu Chen · 2024
Closest in time.
How transformers learn causal structure with gradient descent
Eshaan Nichani, Alex Damian, and Jason D Lee · 2024
Closest in time.
Transformers on markov data: Constant depth suffices
Nived Rajaraman, Marco Bondaschi, Kannan Ramchandran, Michael Gastpar, and Ashok Vardhan Makkuva · 2024
Closest in time.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task
Gautam Reddy · 2024
Closest in time.
Out-of-distribution generalization via composition: a lens through induction heads in transformers
Jiajun Song, Zhuoyan Xu, and Yiqiao Zhong · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Implicit bias of next-token prediction
Christos Thrampoulidis · 2024
Closest in time.
Implicit bias and fast convergence rates for self-attention
Bhavya Vasudeva, Puneesh Deora, and Christos Thrampoulidis · 2024
Closest in time.
Understanding the expressive power and mechanisms of transformer for sequence modeling
Mingze Wang and Weinan E · 2024
Closest in time.
Transformers provably learn sparse token selection while fully-connected nets cannot
Zixuan Wang, Stanley Wei, Daniel Hsu, and Jason D Lee · 2024
Closest in time.