Fetching the paper…
Reading the bibliography…
We show that Transformer encoder architectures can be sped up, with limited accuracy costs, by replacing the self-attention sublayers with simple linear transformations that "mix" input tokens.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Well-read students learn better: On the importance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 1908
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Noam Shazeer. 2019 · 1911
Earlier work this paper cites.
An algorithm for the machine calculation of complex fourier series
James W Cooley and John W Tukey. 1965 · 1965
Earlier work this paper cites.
On the equivalence between one-dimensional discrete walsh-hadamard and multidimensional discrete fourier transforms
Henry O. Kunz. 1979 · 1979
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko. 1989 · 1989
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron. 1993 · 1993
Earlier work this paper cites.
Using fourier-neural recurrent networks to fit sequential input/output data
Renée Koplon and Eduardo D Sontag. 1997 · 1997
Earlier work this paper cites.
Real-time discrimination of ventricular tachyarrhythmia with fourier-transform neural network
Kei-ichiro Minami, Hiroshi Nakajima, and Takeshi Toyoshima. 1999 · 1999
Earlier work this paper cites.
Forenet: fourier recurrent networks for time series prediction
Y Zhang and Lai-Wan Chan. 2000 · 2000
Earlier work this paper cites.
Acceleration of convolutional neural network using fft-based split convolutions
Kamran Chitsaz, Mohsen Hajabdollahi, Nader Karimi, Shadrokh Samavi, and Shahram Shirani. 2020 · 2003
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Fast object/face detection using neural networks and fast fourier transform
Hazem M El-Bakry and Qiangfu Zhao. 2004 · 2004
Earlier work this paper cites.
The design and implementation of fftw3
Matteo Frigo and Steven G Johnson. 2005 · 2005
Earlier work this paper cites.
Synthesizer: Rethinking self-attention in transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. 2020a · 2005
Earlier work this paper cites.
Masked language modeling for proteins via linearly scalable long-context transformers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Jared Davis, Tamas Sarlos, David Belanger, Lucy Colwell, and Adrian Weller. 2020 · 2006
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma. 2020 · 2006
Earlier work this paper cites.
Data movement is all you need: A case study of transformer networks
Andrei Ivanov, Nikoli Dryden, Tal Ben-Nun, Shigang Li, and Torsten Hoefler. 2020 · 2007
Earlier work this paper cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020c · 2009
Earlier work this paper cites.
Pre-trained summarization distillation
Sam Shleifer and Alexander M Rush. 2020 · 2010
Earlier work this paper cites.
Cardiac arrhythmias detection in an ecg beat signal using fast fourier transform and artificial neural network
Himanshu Gothwal, Silky Kedawat, Rajesh Kumar, et al. 2011 · 2011
Earlier work this paper cites.
Rethinking fun: Frequency-domain utilization networks
Kfir Goldberg, Stav Shapiro, Elad Richardson, and Shai Avidan. 2020 · 2012
Cited alongside, same era.
Fast training of convolutional networks through ffts: International conference on learning representations (iclr2014), cbls, april 2014
Michael Mathieu, Mikael Henaff, and Yann LeCun. 2014 · 2014
Cited alongside, same era.
An exploration of parameter redundancy in deep networks with circulant projections
Yu Cheng, Felix X. Yu, Rogerio S. Feris, Sanjiv Kumar, Alok Choudhary, and Shi-Fu Chang. 2015 · 2015
Cited alongside, same era.
Very efficient training of convolutional neural networks using fast fourier transform and overlap-and-add
Tyler Highlander and Andres Rodriguez. 2015 · 2015
Cited alongside, same era.
Fast fourier transform for feature extraction and neural network for classification of electrocardiogram signals
Martina Mironovova and Jirí Bíla. 2015 · 2015
TinyBERT: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020 · 2020
Later among the works it cites.
FastFormers: Highly efficient transformer models for natural language understanding
Young Jin Kim and Hany Hassan. 2020 · 2020
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Later among the works it cites.
Blockwise self-attention for long document understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Wen-tau Yih, Sinong Wang, and Jie Tang. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Structured transforms for small-footprint deep learning
Vikas Sindhwani, Tara N Sainath, and Sanjiv Kumar. 2015 · 2015
Cited alongside, same era.
ACDC: A structured efficient linear layer
Marcin Moczulski, Misha Denil, Jeremy Appleyard, and Nando de Freitas. 2016 · 2016
Cited alongside, same era.
Fcnn: Fourier convolutional neural networks
Harry Pratt, Bryan Williams, Frans Coenen, and Yalin Zheng. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Fft-based deep learning deployment in embedded systems
Sheng Lin, Ning Liu, Mahdi Nazemi, Hongjia Li, Caiwen Ding, Yanzhi Wang, and Massoud Pedram. 2018 · 2018
Cited alongside, same era.
Generating wikipedia by summarizing long sequences
Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018 · 2018
Cited alongside, same era.
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Later among the works it cites.
Fixed encoder self-attention patterns in transformer-based machine translation
Alessandro Raganato, Yves Scherrer, and Jörg Tiedemann. 2020 · 2020
Later among the works it cites.
Language through a prism: A spectral approach for multiscale language representations
Alex Tamkin, Dan Jurafsky, and Noah Goodman. 2020 · 2020
Later among the works it cites.
Fast transformers with clustered attention
Apoorv Vyas, Angelos Katharopoulos, and François Fleuret. 2020 · 2020
Later among the works it cites.
Hard-coded Gaussian attention for neural machine translation
Weiqiu You, Simeng Sun, and Mohit Iyyer. 2020 · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. 2020 · 2020
Later among the works it cites.
A note on more efficient architectures for nlp
Arturs Backurs, Mingda Chen, and Kevin Gimpel. 2021 · 2021
Closest in time.
Rethinking attention with performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamás Sarlós, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J. Colwell, and Adrian Weller. 2021 · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Closest in time.
Longt5: Efficient text-to-text transformer for long sequences
Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontanon, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang. 2021 · 2021
Closest in time.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2021 · 2021
Closest in time.
Fourier neural operator for parametric partial differential equations
Zongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew M. Stuart, and Anima Anandkumar. 2021 · 2021
Closest in time.
Do transformer modifications transfer across implementations and applications?
Sharan Narang, Hyung Won Chung, Yi Tay, Liam Fedus, Thibault Fevry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, and Colin Raffel. 2021 · 2021
Closest in time.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A. Smith, and Lingpeng Kong. 2021 · 2021
Closest in time.
Hopfield networks is all you need
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Lukas Gruber, Markus Holzleitner, Thomas Adler, David P. Kreil, Michael K. Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. 2021 · 2021
Closest in time.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2021 · 2021
Closest in time.
Long range arena : A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. 2021a · 2021
Closest in time.
Mlp-mixer: An all-mlp architecture for vision
Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al. 2021 · 2021
Closest in time.