Fetching the paper…
Reading the bibliography…
Multi-head attention empowers the recent success of transformers, the state-of-the-art models that have achieved remarkable success in sequence modeling and beyond.
The Fourier Integral and Certain of its Applications
N. Wiener · 1933
Earlier work this paper cites.
Remarks on some nonparametric estimates of a density function
M. Rosenblatt · 1956
Earlier work this paper cites.
Lectures on Fourier Integrals
S. Bochner · 1959
Earlier work this paper cites.
On estimation of a probability density function and mode
E. Parzen · 1962
Earlier work this paper cites.
On estimating regression
E. Nadaraya · 1964
Earlier work this paper cites.
Mean square error properties of density estimates
K. B. Davis · 1975
Earlier work this paper cites.
On the optimal rates of convergence for nonparametric deconvolution problems
J. Fan · 1991
Earlier work this paper cites.
Error analysis for general multivariate kernel estimators
M. Wand · 1992
Earlier work this paper cites.
Kernel estimators for multivariate regression
J. Staniswalis, K. Messer, and D. Finston · 1993
Earlier work this paper cites.
Comparison of smoothing parameterizations in bivariate kernel density estimation
M. Wand and M. Jones · 1993
Earlier work this paper cites.
An introduction to the finite element method
J. Reddy · 2004
Earlier work this paper cites.
All of Nonparametric Statistics
L. Wasserman · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Introduction to Nonparametric Estimation
A. Tsybakov · 2009
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
A decomposable attention model for natural language inference
A. Parikh, O. Täckström, D. Das, and J. Uszkoreit · 2016
Earlier work this paper cites.
Parseval networks: Improving robustness to adversarial examples
M. Cisse, P. Bojanowski, E. Grave, Y. Dauphin, and N. Usunier · 2017
Earlier work this paper cites.
Y. Kim, C. Denton, L. Hoang, and A. M. Rush · 2017
Earlier work this paper cites.
A structured self-attentive sentence embedding
Z. Lin, M. Feng, C. N. dos Santos, M. Yu, B. Xiang, B. Zhou, and Y. Bengio · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Multivariate Kernel Smoothing and its Applications
J. Chacón and T. Duong · 2018
Earlier work this paper cites.
Spectral normalization for generative adversarial networks
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida · 2018
Earlier work this paper cites.
All-but-the-top: Simple and effective postprocessing for word representations
J. Mu and P. Viswanath · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Earlier work this paper cites.
Self-attention with relative position representations
P. Shaw, J. Uszkoreit, and A. Vaswani · 2018
Earlier work this paper cites.
Lipschitz-margin training: Scalable certification of perturbation invariance for deep neural networks
Y. Tsuzuku, I. Sato, and M. Sugiyama · 2018
Earlier work this paper cites.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
Cited alongside, same era.
Character-level language modeling with deeper self-attention
R. Al-Rfou, D. Choe, N. Constant, M. Guo, and L. Jones · 2019
Cited alongside, same era.
Sorting out lipschitz function approximation
C. Anil, J. Lucas, and R. Grosse · 2019
Cited alongside, same era.
Adaptive input representations for neural language modeling
A. Baevski and M. Auli · 2019
Cited alongside, same era.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 2019
Cited alongside, same era.
What does BERT look at? an analysis of BERT’s attention
K. Clark, U. Khandelwal, O. Levy, and C. D. Manning · 2019
Cited alongside, same era.
Poor man’s bert: Smaller and faster transformer models
H. Sajjad, F. Dalvi, N. Durrani, and P. Nakov · 2020
Later among the works it cites.
Efficient transformers: A survey
Y. Tay, M. Dehghani, D. Bahri, and D. Metzler · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
S. Wang, B. Li, M. Khabsa, H. Fang, and H. Ma · 2020
Later among the works it cites.
Vivit: A video vision transformer
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid · 2021
Later among the works it cites.
Choose a transformer: Fourier or galerkin
S. Cao · 2021
Later among the works it cites.
Rethinking attention with performers
K. M. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Transformer-xl: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. V. Le, and R. Salakhutdinov · 2019
Cited alongside, same era.
Universal transformers
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
K. Ethayarajh · 2019
Cited alongside, same era.
Designing and interpreting probes with control tasks
J. Hewitt and P. Liang · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Cited alongside, same era.
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Later among the works it cites.
Multiscale vision transformers
H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer · 2021
Later among the works it cites.
Probabilistic attention for interactive segmentation
P. Gabbur, M. Bilkhu, and J. Movellan · 2021
Later among the works it cites.
Pct: Point cloud transformer
M.-H. Guo, J.-X. Cai, Z.-N. Liu, T.-J. Mu, R. R. Martin, and S.-M. Hu · 2021
Later among the works it cites.
Multivariate smoothing via the Fourier integral theorem and Fourier kernel
N. Ho and S. Walker · 2021
Later among the works it cites.
Highly accurate protein structure prediction with alphafold
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko, et al · 2021
Later among the works it cites.
Transformers in vision: A survey
S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah · 2021
Later among the works it cites.
The lipschitz constant of self-attention
H. Kim, G. Papamakarios, and A. Mnih · 2021
Later among the works it cites.
Rethinking graph transformers with spectral attention
D. Kreuzer, D. Beaini, W. Hamilton, V. Létourneau, and P. Tossou · 2021
Later among the works it cites.
T. Lin, Y. Wang, X. Liu, and X. Qiu · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Later among the works it cites.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
A. Rives, J. Meier, T. Sercu, S. Goyal, Z. Lin, J. Liu, D. Guo, M. Ott, C. L. Zitnick, J. Ma, et al · 2021
Later among the works it cites.
Linear transformers are secretly fast weight programmers
I. Schlag, K. Irie, and J. Schmidhuber · 2021
Later among the works it cites.
Probabilistic transformer for time series analysis
B. Tang and D. S. Matteson · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2021
Later among the works it cites.
Modeling concentrated cross-attention for neural machine translation with Gaussian mixture model
S. Zhang and Y. Feng · 2021
Later among the works it cites.
Point transformer
H. Zhao, L. Jiang, J. Jia, P. H. Torr, and V. Koltun · 2021
Later among the works it cites.
Video swin transformer
Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu · 2022
Closest in time.
Improving transformers with probabilistic attention keys
T. Nguyen, T. Nguyen, D. Le, K. Nguyen, A. Tran, R. Baraniuk, N. Ho, and S. Osher · 2022
Closest in time.
Sinkformers: Transformers with doubly stochastic attention
M. E. Sander, P. Ablin, M. Blondel, and G. Peyré · 2022
Closest in time.