Fetching the paper…
Reading the bibliography…
The Transformer architecture is widely applied in sequence modeling applications, yet the theoretical understanding of its working principles remains limited.
Are Transformers universal approximators of sequence-to-sequence functions?
Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S. J., and Kumar, S · 1912
Earlier work this paper cites.
The Theory of Approximation
Jackson, D · 1930
Earlier work this paper cites.
Dimension of metric spaces and Hilbert’s problem 13
Ostrand, P. A · 1936
Earlier work this paper cites.
The Proper Orthogonal Decomposition in the Analysis of Turbulent Flows
Berkooz, G., Holmes, P., and Lumley, J. L · 1993
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Barron, A. R · 1994
Earlier work this paper cites.
Nonlinear approximation
DeVore, R. A · 1998
Earlier work this paper cites.
Spectral properties of the kernel matrix and their relation to kernel methods in machine learning
Braun, M · 2005
Earlier work this paper cites.
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Approximation by Combinations of ReLU and Squared ReLU Ridge Functions with $ \ell^1 $ and $ \ell^0 $ Controls, May 2018
Klusowski, J. M. and Barron, A. R · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Low-Rank Bottleneck in Multi-head Attention Models, February 2020
Bhojanapalli, S., Yun, C., Rawat, A. S., Reddi, S. J., and Kumar, S · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners, July 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
On the Relationship between Self-Attention and Convolutional Layers
Cordonnier, J.-B., Loukas, A., and Jaggi, M · 2020
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2020
Cited alongside, same era.
Limits to Depth Efficiencies of Self-Attention
Levine, Y., Wies, N., Sharir, O., Bata, H., and Shashua, A · 2020
Cited alongside, same era.
On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization Analysis
Universal Approximation Under Constraints is Possible with Transformers, February 2022
Kratsios, A., Zamanlooy, B., Liu, T., and Dokmanić, I · 2022
Later among the works it cites.
Approximation and Optimization Theory for Linear Continuous-Time Recurrent Neural Networks
Li, Z., Han, J., E, W., and Li, Q · 2022
Later among the works it cites.
Your Transformer May Not be as Powerful as You Expect, October 2022
Luo, S., Li, S., Zheng, S., Liu, T.-Y., Wang, L., and He, D · 2022
Later among the works it cites.
Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting, January 2022
Wu, H., Xu, J., Wang, J., and Long, M · 2022
Later among the works it cites.
Are Transformers Effective for Time Series Forecasting?, August 2022
Zeng, A., Chen, M., Zhang, L., and Xu, Q · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, Z., Han, J., E, W., and Li, Q · 2020
Cited alongside, same era.
Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth
Dong, Y., Cordonnier, J.-B., and Loukas, A · 2021
Cited alongside, same era.
On the rate of convergence of a classifier based on a Transformer encoder, November 2021
Gurevych, I., Kohler, M., and Sahin, G. G · 2021
Cited alongside, same era.
Approximation Theory of Convolutional Architectures for Time Series Modelling
Jiang, H., Li, Z., and Li, Q · 2021
Cited alongside, same era.
Optimal Approximation Rate of ReLU Networks in terms of Width and Depth
Shen, Z., Yang, H., and Zhang, S · 2021
Cited alongside, same era.
Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting, March 2021
Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W · 2021
Cited alongside, same era.
Inductive Biases and Variable Creation in Self-Attention Mechanisms, June 2022
Edelman, B. L., Goel, S., Kakade, S., and Zhang, C · 2022
Cited alongside, same era.
Can Vision Transformers Perform Convolution?
Li, S., Chen, X., He, D., and Hsieh, C.-J
Cited in the paper.
Bai, Y., Chen, F., Wang, H., Xiong, C., and Mei, S · 2023
Closest in time.
Looped Transformers as Programmable Computers
Giannou, A., Rajput, S., Sohn, J.-Y., Lee, K., Lee, J. D., and Papailiopoulos, D · 2023
Closest in time.
Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?, July 2023
Kajitsuka, T. and Sato, I · 2023
Closest in time.
Representational Strengths and Limitations of Transformers, June 2023
Sanford, C., Hsu, D., and Telgarsky, M · 2023
Closest in time.
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input, May 2023
Takakura, S. and Suzuki, T · 2023
Closest in time.
Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling, July 2024
Wang, M. and E, W · 2024
Closest in time.