Fetching the paper…
Reading the bibliography…
Recent advances in Transformer architectures have empowered their empirical success in a variety of tasks across different domains.
Remarks on some nonparametric estimates of a density function
M. Rosenblatt · 1956
Earlier work this paper cites.
On estimation of a probability density function and mode
E. Parzen · 1962
Earlier work this paper cites.
On estimating regression
E. A. Nadaraya · 1964
Earlier work this paper cites.
Robust non-linear regression using the dogleg algorithm
R. E. Welsch and R. A. Becker · 1975
Earlier work this paper cites.
Robust statistics: the approach based on influence functions
F. R. Hampel, E. M. Ronchetti, P. Rousseeuw, and W. A. Stahel · 1986
Earlier work this paper cites.
Random generation of combinatorial structures from a uniform distribution
M. R. Jerrum, L. G. Valiant, and V. V. Vazirani · 1986
Earlier work this paper cites.
Robust estimation of a location parameter
P. J. Huber · 1992
Earlier work this paper cites.
The space complexity of approximating the frequency moments
N. Alon, Y. Matias, and M. Szegedy · 1996
Earlier work this paper cites.
Robust regression
J. Fox and S. Weisberg · 2002
Earlier work this paper cites.
An introduction to the finite element method
J. Reddy · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Robust statistics
P. J. Huber · 2011
Earlier work this paper cites.
Robust kernel density estimation
J. Kim and C. Scott · 2012
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2014
Earlier work this paper cites.
Robust kernel density estimation by scaling and projection in hilbert space
R. A. Vandermeulen and C. Scott · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
The uea multivariate time series classification archive, 2018
A. Bagnall, H. A. Dau, J. Lines, M. Flynn, J. Large, A. Bostrom, P. Southam, and E. Keogh · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Earlier work this paper cites.
Adversarial risk and the dangers of evaluating against weak attacks
J. Uesato, B. O’donoghue, P. Kohli, and A. Oord · 2018
Earlier work this paper cites.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
Earlier work this paper cites.
Character-level language modeling with deeper self-attention
R. Al-Rfou, D. Choe, N. Constant, M. Guo, and L. Jones · 2019
Earlier work this paper cites.
Adaptive input representations for neural language modeling
A. Baevski and M. Auli · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. V. Le, and R. Salakhutdinov · 2019
Cited alongside, same era.
Universal transformers
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Benchmarking neural network robustness to common corruptions and perturbations
D. Hendrycks and T. Dietterich · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Cited alongside, same era.
Offline reinforcement learning as one big sequence modeling problem
M. Janner, Q. Li, and S. Levine · 2021
Later among the works it cites.
Transformers in vision: A survey
S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah · 2021
Later among the works it cites.
Rethinking graph transformers with spectral attention
D. Kreuzer, D. Beaini, W. Hamilton, V. Létourneau, and P. Tossou · 2021
Later among the works it cites.
T. Lin, Y. Wang, X. Liu, and X. Qiu · 2021
Later among the works it cites.
Crisisbert: a robust transformer for crisis classification and contextual crisis embedding
J. Liu, T. Singhal, L. T. Blessing, K. L. Wood, and K. H. Lim · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding and improving transformer from a multi-particle dynamic system point of view
Y. Lu, Z. Li, D. He, Z. Sun, B. Dong, T. Qin, L. Wang, and T.-Y. Liu · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Cited alongside, same era.
Transformer dissection: An unified understanding for transformer’s attention via the lens of kernel
Y.-H. H. Tsai, S. Bai, M. Yamada, L.-P. Morency, and R. Salakhutdinov · 2019
Cited alongside, same era.
Learning robust global representations by penalizing local predictive power
H. Wang, S. Ge, Z. Lipton, and E. P. Xing · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. Salakhutdinov, and Q. V. Le · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
On the robustness of vision transformers to adversarial examples
K. Mahmood, R. Mahmood, and M. Van Dijk · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Later among the works it cites.
Linear transformers are secretly fast weight programmers
I. Schlag, K. Irie, and J. Schmidhuber · 2021
Later among the works it cites.
Probabilistic transformer for time series analysis
B. Tang and D. S. Matteson · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2021
Later among the works it cites.
Modeling concentrated cross-attention for neural machine translation with Gaussian mixture model
S. Zhang and Y. Feng · 2021
Later among the works it cites.
Robust kernel density estimation with median-of-means principle
P. Humbert, B. Le Bars, and L. Minvielle · 2022
Closest in time.
Video swin transformer
Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu · 2022
Closest in time.
Towards robust vision transformer
X. Mao, G. Qi, Y. Chen, X. Li, R. Duan, S. Ye, Y. He, and H. Xue · 2022
Closest in time.
Improving transformer with an admixture of attention heads
T. Nguyen, T. Nguyen, H. Do, K. Nguyen, V. Saragadam, M. Pham, K. Nguyen, N. Ho, and S. Osher · 2022
Closest in time.
Improving transformers with probabilistic attention keys
T. Nguyen, T. Nguyen, D. Le, K. Nguyen, A. Tran, R. Baraniuk, N. Ho, and S. Osher · 2022
Closest in time.
Fourierformer: Transformer meets generalized Fourier integral theorem
T. Nguyen, M. Pham, T. Nguyen, K. Nguyen, S. J. Osher, and N. Ho · 2022
Closest in time.
Vision transformers are robust learners
S. Paul and P.-Y. Chen · 2022
Closest in time.
Sinkformers: Transformers with doubly stochastic attention
M. E. Sander, P. Ablin, M. Blondel, and G. Peyré · 2022
Closest in time.
Backdoor attacks on vision transformers
A. Subramanya, A. Saha, S. A. Koohpayegani, A. Tejankar, and H. Pirsiavash · 2022
Closest in time.
Flowformer: Linearizing transformers with conservation flows
H. Wu, J. Wu, J. Xu, J. Wang, and M. Long · 2022
Closest in time.
Tableformer: Robust transformer modeling for table-text encoding
J. Yang, A. Gupta, S. Upadhyay, L. He, R. Goel, and S. Paul · 2022
Closest in time.
Understanding the robustness in vision transformers
D. Zhou, Z. Yu, E. Xie, C. Xiao, A. Anandkumar, J. Feng, and J. M. Alvarez · 2022
Closest in time.