Fetching the paper…
Reading the bibliography…
Many neural network architectures are known to be Turing Complete, and can thus, in principle implement arbitrary algorithms.
A generalized representer theorem
Bernhard Schölkopf, Ralf Herbrich, and Alex J Smola · 2001
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Transformer dissection: a unified understanding of transformer’s attention via the lens of kernel
Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Martin J Wainwright · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Earlier work this paper cites.
Xcit: Cross-covariance image transformers
Alaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski, Matthijs Douze, Armand Joulin, Ivan Laptev, Natalia Neverova, Gabriel Synnaeve, Jakob Verbeek, et al · 2021
Earlier work this paper cites.
Attention is turing complete
Jorge Pérez, Pablo Barceló, and Javier Marinkovic · 2021
Earlier work this paper cites.
Linear transformers are secretly fast weight programmers
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber · 2021
Earlier work this paper cites.
Transformers are deep infinite-dimensional non-mercer binary kernel machines
Matthew A Wright and Joseph E Gonzalez · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2022
Cited alongside, same era.
Kerple: Kernelized relative positional embedding for length extrapolation
Ta-Chung Chi, Ting-Han Fan, Peter J Ramadge, and Alexander Rudnicky · 2022
Cited alongside, same era.
Why can gpt learn in-context? language models implicitly perform gradient descent as meta-optimizers
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant · 2022
Primal-attention: Self-attention through asymmetric kernel svd in primal representation
Yingyi Chen, Qinghua Tao, Francesco Tonin, and Johan AK Suykens · 2023
Closest in time.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D Lee, and Dimitris Papailiopoulos · 2023
Closest in time.
In-context convergence of transformers
Yu Huang, Yuan Cheng, and Yingbin Liang · 2023
Closest in time.
Licong Lin, Yu Bai, and Song Mei · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Cited alongside, same era.
Statistically meaningful approximation: a case study on approximating turing machines with transformers
Colin Wei, Yining Chen, and Tengyu Ma · 2022
Cited alongside, same era.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra · 2023
Cited alongside, same era.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2023
Cited alongside, same era.
Improving transformers with probabilistic attention keys
Tam Minh Nguyen, Tan Minh Nguyen, Dung DD Le, Duy Khuong Nguyen, Viet-Anh Tran, Richard Baraniuk, Nhat Ho, and Stanley Osher
Cited in the paper.
Fourierformer: Transformer meets generalized fourier integral theorem
Tan Nguyen, Minh Pham, Tam Nguyen, Khai Nguyen, Stanley Osher, and Nhat Ho
Cited in the paper.
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov
Cited in the paper.
Arvind Mahankali, Tatsunori B Hashimoto, and Tengyu Ma · 2023
Closest in time.
Replacing softmax with relu in vision transformers
Mitchell Wortsman, Jaehoon Lee, Justin Gilmer, and Simon Kornblith · 2023
Closest in time.
How many pretraining tasks are needed for in-context learning of linear regression?
Jingfeng Wu, Difan Zou, Zixiang Chen, Vladimir Braverman, Quanquan Gu, and Peter L Bartlett · 2023
Closest in time.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L Bartlett · 2023
Closest in time.