Fetching the paper…
Reading the bibliography…
Despite the widespread adoption of prompting, prompt tuning and prefix-tuning of transformer models, our theoretical understanding of these fine-tuning methods remains limited.
Beiträge zur Theorie der Kugelfunktionen
Paul Funk. 1915 · 1915
Earlier work this paper cites.
Über orthogonal-invariante Integralgleichungen
E Hecke. 1917 · 1917
Earlier work this paper cites.
On the representation of continuous functions of many variables by superposition of continuous functions of one variable and addition
Andrei Nikolaevich Kolmogorov. 1957 · 1957
Earlier work this paper cites.
Covering a sphere with spheres
C. A. Rogers. 1963 · 1963
Earlier work this paper cites.
Constructive polynomial approximation on spheres and projective spaces
David L Ragozin. 1971 · 1971
Earlier work this paper cites.
Computation of modified Bessel functions and their ratios
Donald E Amos. 1974 · 1974
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko. 1989 · 1989
Earlier work this paper cites.
Representation properties of networks: Kolmogorov’s theorem is irrelevant
Federico Girosi and Tomaso Poggio. 1989 · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White. 1989 · 1989
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron. 1993 · 1993
Earlier work this paper cites.
Approximation by spherical convolution
Valdir Antônio Menegatto. 1997 · 1997
Earlier work this paper cites.
On the optimality of the random hyperplane rounding technique for MAX CUT
Uriel Feige and Gideon Schechtman. 2002 · 2002
Earlier work this paper cites.
Concise formulas for the area and volume of a hyperspherical cap
Shengqiao Li. 2010 · 2010
Earlier work this paper cites.
Spherical Harmonics and Approximations on the Unit Sphere: An Introduction
Kendall Atkinson and Weimin Han. 2012 · 2012
Earlier work this paper cites.
Approximation Theory and Harmonic Analysis on Spheres and Balls
Feng Dai and Yuan Xu. 2013 · 2013
Earlier work this paper cites.
On radial functions and distributions and their Fourier transforms
Ricardo Estrada. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Representation benefits of deep feedforward networks
Matus Telgarsky. 2015 · 2015
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Certain inequalities of Kober and Lazarević type
Yogesh J Bagul and Satish K Panchal. 2018 · 2018
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Cited alongside, same era.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar. 2019 · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2020
Cited alongside, same era.
Stolen probability: A structural weakness of neural language models
David Demeter, Gregory Kimmel, and Doug Downey. 2020 · 2020
Cited alongside, same era.
Sumformer: Universal approximation for efficient transformers
Silas Alberti, Niclas Dern, Laura Thesing, and Gitta Kutyniok. 2023 · 2023
Later among the works it cites.
Soft prompting might be a bug, not a feature
Luke Bailey, Gustaf Ahdritz, Anat Kleiman, Siddharth Swaroop, Finale Doshi-Velez, and Weiwei Pan. 2023 · 2023
Later among the works it cites.
On the expressivity role of LayerNorm in transformers’ attention
Shaked Brody, Uri Alon, and Eran Yahav. 2023 · 2023
Later among the works it cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. 2023 · 2023
Later among the works it cites.
Perfectly secure steganography using minimum entropy coupling
Christian Schroeder de Witt, Samuel Sokota, J. Zico Kolter, Jakob Foerster, and Martin Strohmeier. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
AutoPrompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Cited alongside, same era.
Lower bounds on the modified Bessel function of the first kind
Alex Barnett. 2021 · 2021
Cited alongside, same era.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. 2021 · 2021
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Cited alongside, same era.
Prefix-Tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Cited alongside, same era.
On the expressive power of self-attention matrices
Valerii Likhosherstov, Krzysztof Choromanski, and Adrian Weller. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
On the optimization and generalization of multi-head attention
Puneesh Deora, Rouzbeh Ghaderi, Hossein Taheri, and Christos Thrampoulidis. 2023 · 2023
Later among the works it cites.
LLM-Adapters: An adapter family for parameter-efficient fine-tuning of large language models
Zhiqiang Hu, Yihuai Lan, Lei Wang, Wanyu Xu, Ee-Peng Lim, Roy Ka-Wei Lee, Lidong Bing, and Soujanya Poria. 2023 · 2023
Later among the works it cites.
Approximation theory of transformer networks for sequence modeling
Haotian Jiang and Qianxiao Li. 2023 · 2023
Later among the works it cites.
A brief survey on the approximation theory for sequence modelling
Haotian Jiang, Qianxiao Li, Zhong Li, and Shida Wang. 2023 · 2023
Later among the works it cites.
Scaling down to scale up: A guide to parameter-efficient fine-tuning
Vladislav Lialin, Vijeta Deshpande, and Anna Rumshisky. 2023 · 2023
Later among the works it cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023 · 2023
Later among the works it cites.
Memorization capacity of multi-head attention in transformers
Sadegh Mahdavi, Renjie Liao, and Christos Thrampoulidis. 2023 · 2023
Later among the works it cites.
Representational strengths and limitations of transformers
Clayton Sanford, Daniel Hsu, and Matus Telgarsky. 2023 · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, Joao Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. 2023 · 2023
Later among the works it cites.
Universality and limitations of prompt tuning
Yihan Wang, Jatin Chauhan, Wei Wang, and Cho-Jui Hsieh. 2023 · 2023
Later among the works it cites.
Fundamental limitations of alignment in large language models
Yotam Wolf, Noam Wies, Oshri Avnery, Yoav Levine, and Amnon Shashua. 2023 · 2023
Later among the works it cites.
Pretraining data mixtures enable narrow model selection capabilities in transformer models
Steve Yadlowsky, Lyric Doshi, and Nilesh Tripuraneni. 2023 · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Later among the works it cites.
When do prompting and prefix-tuning work? A theory of capabilities and limitations
Aleksandar Petrov, Philip HS Torr, and Adel Bibi. 2024 · 2024
Closest in time.