Fetching the paper…
Reading the bibliography…
Recent advances in efficient Transformers have exploited either the sparsity or low-rank properties of attention matrices to reduce the computational and memory bottlenecks of modeling long sequences.
Analysis of a complex of statistical variables into principal components
Harold Hotelling · 1933
Earlier work this paper cites.
Sparse matrices , volume 69
Reginald P Tewarson and Reginald P Tewarson · 1973
Earlier work this paper cites.
Interpretation of the correlation coefficient: a basic review
Richard Taylor · 1990
Earlier work this paper cites.
Approximate nearest neighbors: towards removing the curse of dimensionality
Piotr Indyk and Rajeev Motwani · 1998
Earlier work this paper cites.
Similarity search in high dimensions via hashing
Aristides Gionis, Piotr Indyk, Rajeev Motwani, et al · 1999
Earlier work this paper cites.
Bravais-pearson and spearman correlation coefficients: meaning, test of hypothesis and confidence interval
R Artusi, P Verderio, and E Marubini · 2002
Earlier work this paper cites.
Exact matrix completion via convex optimization
Emmanuel J Candès and Benjamin Recht · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2009
Earlier work this paper cites.
Robust principal component analysis?
Emmanuel J Candès, Xiaodong Li, Yi Ma, and John Wright · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts · 2011
Earlier work this paper cites.
A simpler approach to matrix completion
Benjamin Recht · 2011
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2011
Earlier work this paper cites.
The acl anthology network corpus
Dragomir R Radev, Pradeep Muthukrishnan, Vahed Qazvinian, and Amjad Abu-Jbara · 2013
Earlier work this paper cites.
Asymmetric lsh (alsh) for sublinear time maximum inner product search (mips)
Anshumali Shrivastava and Ping Li · 2014
Earlier work this paper cites.
Global convergence of stochastic gradient descent for some non-convex matrix problems
Christopher De Sa, Christopher Re, and Kunle Olukotun · 2015
Earlier work this paper cites.
Structured transforms for small-footprint deep learning
Vikas Sindhwani, Tara N. Sainath, and Sanjiv Kumar · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
GPU kernels for block-sparse weights
Scott Gray, Alec Radford, and Diederik P Kingma · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Densified winner take all (wta) hashing for sparse datasets
Beidi Chen and Anshumali Shrivastava · 2018
Earlier work this paper cites.
Unique entity estimation with application to the syrian conflict
Beidi Chen, Anshumali Shrivastava, and Rebecca C Steorts · 2018
Earlier work this paper cites.
A two-pronged progress in structured dense matrix vector multiplication
Christopher De Sa, Albert Gu, Rohan Puttagunta, Christopher Ré, and Atri Rudra · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Learning long-range spatial dependencies with horizontal gated-recurrent units
Drew Linsley, Junkyung Kim, Vijay Veerabadran, and Thomas Serre · 2018
Cited alongside, same era.
Listops: A diagnostic dataset for latent tree learning
Nikita Nangia and Samuel R Bowman · 2018
Cited alongside, same era.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Cited alongside, same era.
Kaleidoscope: An efficient, learnable representation for all structured linear maps
Tri Dao, Nimit Sohoni, Albert Gu, Matthew Eichhorn, Amit Blonder, Megan Leszczynski, Atri Rudra, and Christopher Ré · 2020
Later among the works it cites.
Smyrf: Efficient attention using asymmetric clustering
Giannis Daras, Nikita Kitaev, Augustus Odena, and Alexandros G Dimakis · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Sparse GPU kernels for deep learning
Trevor Gale, Matei Zaharia, Cliff Young, and Erich Elsen · 2020
Later among the works it cites.
Hippo: Recurrent memory with optimal polynomial projections
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning compressed transforms with low displacement rank
Anna Thomas, Albert Gu, Tri Dao, Atri Rudra, and Christopher Ré · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Cited alongside, same era.
Fast and accurate stochastic gradient estimation
Beidi Chen, Yingchen Xu, and Anshumali Shrivastava · 2019
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Learning fast algorithms for linear transforms using butterfly factorizations
Tri Dao, Albert Gu, Matthew Eichhorn, Atri Rudra, and Christopher Ré · 2019
Cited alongside, same era.
Learning space partitions for nearest neighbor search
Yihe Dong, Piotr Indyk, Ilya Razenshteyn, and Tal Wagner · 2019
Cited alongside, same era.
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Later among the works it cites.
Sub-linear memory: How to make performers slim
Valerii Likhosherstov, Krzysztof Choromanski, Jared Davis, Xingyou Song, and Adrian Weller · 2020
Later among the works it cites.
Climbing the wol: Training for cheaper inference
Zichang Liu, Zhaozhuo Xu, Alan Ji, Jonathan Li, Beidi Chen, and Anshumali Shrivastava · 2020
Later among the works it cites.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap · 2020
Later among the works it cites.
Hopfield networks is all you need
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, et al · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell · 2021
Closest in time.
Mongoose: A learnable lsh framework for efficient neural network training
Beidi Chen, Zichang Liu, Binghui Peng, Zhaozhuo Xu, Jonathan Lingjie Li, Tri Dao, Zhao Song, Anshumali Shrivastava, and Christopher Ré · 2021
Closest in time.
A tale of two efficient and informative negative sampling distributions
Shabnam Daghaghi, Tharun Medini, Nicholas Meisburger, Beidi Chen, Mengnan Zhao, and Anshumali Shrivastava · 2021
Closest in time.
Simplified self-attention for transformer-based end-to-end speech recognition
Haoneng Luo, Shiliang Zhang, Ming Lei, and Lei Xie · 2021
Closest in time.
Luna: Linear unified nested attention
Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma, and Luke Zettlemoyer · 2021
Closest in time.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier · 2021
Closest in time.
Nystromformer: A Nystrom-based algorithm for approximating self-attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis EH Tay, Jiashi Feng, and Shuicheng Yan · 2021
Closest in time.
Long-short transformer: Efficient transformers for language and vision
Chen Zhu, Wei Ping, Chaowei Xiao, Mohammad Shoeybi, Tom Goldstein, Anima Anandkumar, and Bryan Catanzaro · 2021
Closest in time.