Fetching the paper…
Reading the bibliography…
Natural language processing (NLP) made an impressive jump with the introduction of Transformers.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H · 1989
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Barron, A. R · 1994
Earlier work this paper cites.
Neural networks for optimal approximation of smooth and analytic functions
Mhaskar, H. N · 1996
Earlier work this paper cites.
When is the algebra of multisymmetric polynomials generated by the elementary multisymmetric polynomials?
Briand, E · 2004
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
The expressive power of neural networks: A view from the width
Lu, Z., Pu, H., Wang, F., Hu, Z., and Wang, L · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Error bounds for approximations with deep relu networks
Yarotsky, D · 2017
Earlier work this paper cites.
Deep sets
Zaheer, M., Kottur, S., Ravanbakhsh, S., Poczos, B., Salakhutdinov, R. R., and Smola, A. J · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Optimal approximation of piecewise smooth functions using deep relu neural networks
Petersen, P. and Voigtlaender, F · 2018
Earlier work this paper cites.
Provable approximation properties for deep neural networks
Shaham, U., Cloninger, A., and Coifman, R. R · 2018
Earlier work this paper cites.
Optimal approximation with sparsely connected deep neural networks
Bolcskei, H., Grohs, P., Kutyniok, G., and Petersen, P · 2019
Cited alongside, same era.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Cited alongside, same era.
Deep learning: new computational modelling techniques for genomics
Eraslan, G., Avsec, Ž., Gagneur, J., and Theis, F. J · 2019
Cited alongside, same era.
On the limitations of representing functions on sets
Wagstaff, E., Fuchs, F., Engelcke, M., Posner, I., and Osborne, M. A · 2019
Cited alongside, same era.
Are transformers universal approximators of sequence-to-sequence functions?
Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S. J., and Kumar, S · 2019
Cited alongside, same era.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Later among the works it cites.
Long range arena: A benchmark for efficient transformers, 2020
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Later among the works it cites.
o ( n ) o(n) connections are expressive enough: Universal approximability of sparse transformers
Yun, C., Chang, Y.-W., Bhojanapalli, S., Rawat, A. S., Reddi, S., and Kumar, S · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Error bounds for approximations with deep relu neural networks in w s, p norms
Gühring, I., Kutyniok, G., and Petersen, P · 2020
Cited alongside, same era.
On representing (anti) symmetric functions
Hutter, M · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Later among the works it cites.
Highly accurate protein structure prediction with alphafold
Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., et al · 2021
Later among the works it cites.
On the expressive power of self-attention matrices
Likhosherstov, V., Choromanski, K., and Weller, A · 2021
Later among the works it cites.
Representation theorem for multivariable totally symmetric functions
Chen, C., Chen, Z., and Lu, J · 2022
Later among the works it cites.
Efficient transformers: A survey
Tay, Y., Dehghani, M., Bahri, D., and Metzler, D · 2022
Later among the works it cites.
Transformers in time series: A survey
Wen, Q., Zhou, T., Zhang, C., Chen, W., Ma, Z., Yan, J., and Sun, L · 2022
Later among the works it cites.
Universal approximations of invariant maps by neural networks
Yarotsky, D · 2022
Later among the works it cites.