Fetching the paper…
Reading the bibliography…
Transformers are deep architectures that define "in-context mappings" which enable predicting new tokens based on a given set of tokens (such as a prompt in NLP applications or a set of patches for a vision transformer).
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Approximation theory of the mlp model in neural networks
Allan Pinkus · 1999
Earlier work this paper cites.
Infinite dimensional analysis
Charalambos D. Aliprantis and Kim C Border · 2006
Earlier work this paper cites.
Support theorems for the radon transform and cramér-wold theorems
Jan Boman and Filip Lindskog · 2009
Earlier work this paper cites.
Optimal transport for applied mathematicians
Filippo Santambrogio · 2015
Earlier work this paper cites.
Approximating continuous functions by relu nets of minimal width
Boris Hanin and Mark Sellke · 2017
Earlier work this paper cites.
General topology
John L Kelley · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Error bounds for approximations with deep relu networks
Dmitry Yarotsky · 2017
Earlier work this paper cites.
How powerful are graph neural networks?
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka · 2018
Earlier work this paper cites.
Stochastic deep networks
Gwendoline De Bie, Gabriel Peyré, and Marco Cuturi · 2019
Earlier work this paper cites.
Universal invariant and equivariant graph neural networks
Nicolas Keriven and Gabriel Peyré · 2019
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J Reddi, and Sanjiv Kumar · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Deep neural networks, generic universal interpolation, and controlled odes
Christa Cuchiero, Martin Larsson, and Josef Teichmann · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
A mathematical theory of attention
James Vuckovic, Aristide Baratin, and Remi Tachet des Combes · 2020
Cited alongside, same era.
A mathematical perspective on transformers
Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet · 2023
Later among the works it cites.
Neural operator: Learning maps between function spaces with applications to pdes
Nikola Kovachki, Zongyi Li, Burigede Liu, Kamyar Azizzadenesheli, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar · 2023
Later among the works it cites.
Small transformers compute universal metric embeddings
Anastasis Kratsios, Valentin Debarnot, and Ivan Dokmanić · 2023
Later among the works it cites.
An approximation theory for metric space-valued functions with a view towards deep learning
Anastasis Kratsios, Chong Liu, Matti Lassas, Maarten V de Hoop, and Ivan Dokmanić · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ding-Xuan Zhou · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Cited alongside, same era.
Dnn expression rate analysis of high-dimensional pdes: Application to option pricing
Dennis Elbrächter, Philipp Grohs, Arnulf Jentzen, and Christoph Schwab · 2022
Cited alongside, same era.
Universal approximation theorems for differentiable geometric deep learning
Anastasis Kratsios and Léonie Papon · 2022
Cited alongside, same era.
Your transformer may not be as powerful as you expect
Shengjie Luo, Shanda Li, Shuxin Zheng, Tie-Yan Liu, Liwei Wang, and Di He · 2022
Cited alongside, same era.
Sinkformers: Transformers with doubly stochastic attention
Michael E Sander, Pierre Ablin, Mathieu Blondel, and Gabriel Peyré · 2022
Cited alongside, same era.
Universal approximation power of deep residual neural networks through the lens of control
Paulo Tabuada and Bahman Gharesifard · 2022
Cited alongside, same era.
Sumformer: Universal approximation for efficient transformers
Silas Alberti, Niclas Dern, Laura Thesing, and Gitta Kutyniok · 2023
Cited alongside, same era.
Arvind Mahankali, Tatsunori B Hashimoto, and Tengyu Ma · 2023
Later among the works it cites.
The expresssive power of transformers with chain of thought
William Merrill and Ashish Sabharwal · 2023
Later among the works it cites.
Attending to graph transformers
Luis Müller, Mikhail Galkin, Christopher Morris, and Ladislav Rampášek · 2023
Later among the works it cites.
Uncovering mesa-optimization algorithms in transformers
Johannes von Oswald, Eyvind Niklasson, Maximilian Schlegel, Seijin Kobayashi, Nicolas Zucchet, Nino Scherrer, Nolan Miller, Mark Sandler, Max Vladymyrov, Razvan Pascanu, et al · 2023
Later among the works it cites.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L Bartlett · 2023
Later among the works it cites.
Andrei Agrachev and Cyril Letrouit · 2024
Closest in time.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra · 2024
Closest in time.
How smooth is attention?
Valérie Castin, Pierre Ablin, and Gabriel Peyré · 2024
Closest in time.
Transformers are expressive, but are they expressive enough for regression?
Swaroop Nath, Harshad Khadilkar, and Pushpak Bhattacharyya · 2024
Closest in time.
How do transformers perform in-context autoregressive learning?
Michael E Sander, Raja Giryes, Taiji Suzuki, Mathieu Blondel, and Gabriel Peyré · 2024
Closest in time.
What formal languages can transformers express? a survey
Lena Strobl, William Merrill, Gail Weiss, David Chiang, and Dana Angluin · 2024
Closest in time.