Fetching the paper…
Reading the bibliography…
Attention-based mechanisms are widely used in machine learning, most prominently in transformers.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Spherical harmonics in p dimensions
Christopher Frye and Costas J Efthimiou · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Earlier work this paper cites.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
On the equivalence between kernel quadrature rules and random feature expansions
Francis Bach · 2017
Earlier work this paper cites.
Depth separation for neural networks
Amit Daniely · 2017
Earlier work this paper cites.
Depth-width tradeoffs in approximating natural functions with neural networks
Itay Safran and Ohad Shamir · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Distribution-specific hardness of learning neural networks
Ohad Shamir · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science
Roman Vershynin · 2018
Earlier work this paper cites.
Depth-width trade-offs for relu networks via sharkovsky’s theorem
Vaggos Chatziafratis, Sai Ganesh Nagarajan, Ioannis Panageas, and Xiao Wang · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J Reddi, and Sanjiv Kumar · 2019
Earlier work this paper cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Earlier work this paper cites.
Root mean square layer normalization
Biao Zhang and Rico Sennrich · 2019
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Compressing pre-trained language models by matrix decomposition
Matan Ben Noach and Yoav Goldberg · 2020
Cited alongside, same era.
Low-rank bottleneck in multi-head attention models
Srinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Cited alongside, same era.
Approximate is good enough: Probabilistic variants of dimensional and margin complexity
Pritish Kamath, Omar Montasser, and Nathan Srebro · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Linearized two-layers neural networks in high dimension
On the optimization and generalization of multi-head attention
Puneesh Deora, Rouzbeh Ghaderi, Hossein Taheri, and Christos Thrampoulidis · 2023
Later among the works it cites.
What can a single attention layer learn? a study through the random features lens
Hengyu Fu, Tianyu Guo, Yu Bai, and Song Mei · 2023
Later among the works it cites.
Are transformers with one layer self-attention using low-rank weight matrices universal approximators?, 2023
Tokio Kajitsuka and Issei Sato · 2023
Later among the works it cites.
On the expressive flexibility of self-attention matrices
Valerii Likhosherstov, Krzysztof Choromanski, and Adrian Weller · 2023
Later among the works it cites.
LightFormer: Light-weight transformer using SVD-based weight transfer and parameter sharing
Xiuqing Lv, Peng Zhang, Sunzhu Li, Guobing Gan, and Yueheng Sun · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Cited alongside, same era.
Compressing pre-trained language models using progressive low rank decomposition
Habib Hajimolahoseini, Mehdi Rezagholizadeh, Vahid Partovinia, Marzieh Tahaei, Omar Mohamed Awad, and Yang Liu · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2022
Cited alongside, same era.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank · 2022
Cited alongside, same era.
An empirical analysis of compute-optimal large language model training
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Thomas Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karén Simonyan, Erich Elsen, Oriol Vinyals, Jack Rae, and Laurent Sifre · 2022
Cited alongside, same era.
Theodor Misiakiewicz and Andrea Montanari · 2023
Later among the works it cites.
The expresssive power of transformers with chain of thought
William Merrill and Ashish Sabharwal · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer
Yuandong Tian, Yiping Wang, Beidi Chen, and Simon S Du · 2023
Later among the works it cites.
Separations in the representational capabilities of transformers and recurrent architectures, 2024
Satwik Bhattamishra, Michael Hahn, Phil Blunsom, and Varun Kanade · 2024
Closest in time.
Scaling laws for associative memories
Vivien Cabannes, Elvis Dohmatob, and Alberto Bietti · 2024
Closest in time.
Siyu Chen, Heejune Sheen, Tianhao Wang, and Zhuoran Yang · 2024
Closest in time.
Olmo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al · 2024
Closest in time.
Repeat after me: Transformers are better than state space models at copying
Samy Jelassi, David Brandfonbrener, Sham M Kakade, and Eran Malach · 2024
Closest in time.
Memorization capacity of multi-head attention in transformers
Sadegh Mahdavi, Renjie Liao, and Christos Thrampoulidis · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
The truth is in there: Improving reasoning in language models with layer-selective rank reduction
Pratyusha Sharma, Jordan T. Ash, and Dipendra Misra · 2024
Closest in time.
Transformers, parallel computation, and logarithmic depth, 2024
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2024
Closest in time.
Representational strengths and limitations of transformers
Clayton Sanford, Daniel J Hsu, and Matus Telgarsky · 2024
Closest in time.
What Formal Languages Can Transformers Express? A Survey
Lena Strobl, William Merrill, Gail Weiss, David Chiang, and Dana Angluin · 2024
Closest in time.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L. Bartlett · 2024
Closest in time.