Fetching the paper…
Reading the bibliography…
Ever since their conception, Transformers have taken over traditional sequence models in many tasks, such as NLP, image classification, and video/audio processing, for their fast training and superior performance.
An algorithm for the machine calculation of complex fourier series
J. W. Cooley and J. W. Tukey · 1965
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
The design and implementation of fftw3
M. Frigo and S. G. Johnson · 2005
Earlier work this paper cites.
Product quantization for nearest neighbor search
H. Jegou, M. Douze, and C. Schmid · 2010
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
A. Graves, G. Wayne, and I. Danihelka · 2014
Earlier work this paper cites.
Conditional computation in neural networks for faster models
E. Bengio, P.-L. Bacon, J. Pineau, and D. Precup · 2015
Earlier work this paper cites.
Bidirectional lstm-crf models for sequence tagging
Z. Huang, W. Xu, and K. Yu · 2015
Earlier work this paper cites.
Sequence to sequence-video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko · 2015
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
P. Helber, B. Bischke, A. Dengel, and D. Borth · 2019
Earlier work this paper cites.
Axial attention in multidimensional transformers
J. Ho, N. Kalchbrenner, D. Weissenborn, and T. Salimans · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Transformer dissection: A unified understanding of transformer’s attention via the lens of kernel
Y.-H. H. Tsai, S. Bai, M. Yamada, L.-P. Morency, and R. Salakhutdinov · 2019
Earlier work this paper cites.
X. Zhu, J. Hu, C. Qiu, Y. Shi, H. Bagheri, J. Kang, H. Li, L. Mou, G. Zhang, M. Häberle, S. Han, Y. Hua, R. Huang, L. Hughes, Y. Sun, M. Schmitt, and Y. Wang · 2019
Cited alongside, same era.
Etc: Encoding long and structured inputs in transformers
J. Ainslie, S. Ontanon, C. Alberti, V. Cvicek, Z. Fisher, P. Pham, A. Ravula, S. Sanghai, Q. Wang, and L. Yang · 2020
Cited alongside, same era.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Transformer in transformer
K. Han, A. Xiao, E. Wu, J. Guo, C. Xu, and Y. Wang · 2021
Later among the works it cites.
Costa: Communication-optimal shuffle and transpose algorithm with process relabeling
M. Kabić, S. Pintarelli, A. Kozhevnikov, and J. VandeVondele · 2021
Later among the works it cites.
Fnet: Mixing tokens with fourier transforms
J. Lee-Thorp, J. Ainslie, I. Eckstein, and S. Ontanon · 2021
Later among the works it cites.
Combiner: Full attention transformer with sparse computation cost
H. Ren, H. Dai, Z. Dai, M. Yang, J. Leskovec, D. Schuurmans, and B. Dai · 2021
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
A. Roy, M. Saffar, A. Vaswani, and D. Grangier · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Davis, A. Mohiuddin, L. Kaiser, et al · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Hippo: Recurrent memory with optimal polynomial projections
A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré · 2020
Cited alongside, same era.
Reformer: The efficient transformer
N. Kitaev, Ł. Kaiser, and A. Levskaya · 2020
Cited alongside, same era.
Hj egou,“training data-efficient image transformers & distillation through attention,”
H. Touvron, M. Cord, M. Douze, F. Massa, and A. Sablay-rolles · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma · 2020
Cited alongside, same era.
Big bird: Transformers for longer sequences
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang, et al · 2020
Cited alongside, same era.
Segatron: Segment-aware transformer for language modeling and understanding
H. Bai, P. Shi, J. Lin, Y. Xie, L. Tan, K. Xiong, W. Gao, and M. Li · 2021
Cited alongside, same era.
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit, et al · 2021
Later among the works it cites.
Rethinking and improving relative position encoding for vision transformer
K. Wu, H. Peng, M. Chen, J. Fu, and H. Chao · 2021
Later among the works it cites.
Positional encoding as spatial inductive bias in gans
R. Xu, X. Wang, K. Chen, B. Zhou, and C. C. Loy · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Later among the works it cites.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Later among the works it cites.
What makes convolutional models great on long sequence modeling?, 2022
Y. Li, T. Cai, Y. Zhang, D. Chen, and D. Dey · 2022
Later among the works it cites.
Learning to drop out: An adversarial approach to training sequence vaes
Đ. Miladinović, K. Shridhar, K. Jain, M. B. Paulus, J. M. Buhmann, and C. Allen · 2022
Later among the works it cites.
Y. Wu, M. N. Rabe, D. Hutchins, and C. Szegedy · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing
Y. Zhou, T. Lei, H. Liu, N. Du, Y. Huang, V. Zhao, A. Dai, Z. Chen, Q. Le, and J. Laudon · 2022
Later among the works it cites.