Fetching the paper…
Reading the bibliography…
Transformers are slow and memory-hungry on long sequences, since the time and memory complexity of self-attention are quadratic in sequence length.
The working set model for program behavior
Peter J Denning · 1968
Earlier work this paper cites.
Displacement ranks of matrices and linear equations
Thomas Kailath, Sun-Yuan Kung, and Martin Morf · 1979
Earlier work this paper cites.
The input/output complexity of sorting and related problems
Alok Aggarwal and S Vitter, Jeffrey · 1988
Earlier work this paper cites.
A data locality optimizing algorithm
Michael E Wolf and Monica S Lam · 1991
Earlier work this paper cites.
Random butterfly transformations with applications in computational linear algebra
D Stott Parker · 1995
Earlier work this paper cites.
Data cube: A relational aggregation operator generalizing group-by, cross-tab, and sub-totals
Jim Gray, Surajit Chaudhuri, Adam Bosworth, Andrew Layman, Don Reichart, Murali Venkatrao, Frank Pellow, and Hamid Pirahesh · 1997
Earlier work this paper cites.
On a new class of structured matrices
Y Eidelman and I Gohberg · 1999
Earlier work this paper cites.
An updated set of basic linear algebra subprograms (blas)
L Susan Blackford, Antoine Petitet, Roldan Pozo, Karin Remington, R Clint Whaley, James Demmel, Jack Dongarra, Iain Duff, Sven Hammarling, Greg Henry, et al · 2002
Earlier work this paper cites.
Memory hierarchy design
John Hennessy and David Patterson · 2003
Earlier work this paper cites.
Database management systems , volume 3
Raghu Ramakrishnan, Johannes Gehrke, and Johannes Gehrke · 2003
Earlier work this paper cites.
Optimal space lower bounds for all frequency moments
David P Woodruff · 2004
Earlier work this paper cites.
Parameterized Complexity Theory
Jörg Flum and Martin Grohe · 2006
Earlier work this paper cites.
Evaluating derivatives: principles and techniques of algorithmic differentiation
Andreas Griewank and Andrea Walther · 2008
Earlier work this paper cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2009
Earlier work this paper cites.
Roofline: an insightful visual performance model for multicore architectures
Samuel Williams, Andrew Waterman, and David Patterson · 2009
Earlier work this paper cites.
Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe · 2013
Earlier work this paper cites.
Parallel stochastic gradient algorithms for large-scale matrix completion
Benjamin Recht and Christopher Ré · 2013
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Song Han, Jeff Pool, John Tran, and William J Dally · 2015
Earlier work this paper cites.
Scalability! but at what { \{ COST } \} ?
Frank McSherry, Michael Isard, and Derek G Murray · 2015
Earlier work this paper cites.
Structured transforms for small-footprint deep learning
Vikas Sindhwani, Tara Sainath, and Sanjiv Kumar · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J Dally · 2016
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark · 2016
Earlier work this paper cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Xin Dong, Shangyu Chen, and Sinno Jialin Pan · 2017
Earlier work this paper cites.
In-datacenter performance analysis of a tensor processing unit
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al · 2017
Earlier work this paper cites.
Runtime neural pruning
Ji Lin, Yongming Rao, Jiwen Lu, and Jie Zhou · 2017
Earlier work this paper cites.
Nvidia Tesla V100 GPU architecture, 2017
NVIDIA · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A two-pronged progress in structured dense matrix vector multiplication
Christopher De Sa, Albert Gu, Rohan Puttagunta, Christopher Ré, and Atri Rudra · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
Dissecting the nvidia Volta GPU architecture via microbenchmarking
Zhe Jia, Marco Maggioni, Benjamin Staiger, and Daniele P Scarpazza · 2018
Earlier work this paper cites.
Online normalizer calculation for softmax
Maxim Milakov and Natalia Gimelshein · 2018
Cited alongside, same era.
Neural legal judgment prediction in English
Ilias Chalkidis, Ion Androutsopoulos, and Nikolaos Aletras · 2019
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Learning fast algorithms for linear transforms using butterfly factorizations
Tri Dao, Albert Gu, Matthew Eichhorn, Atri Rudra, and Christopher Ré · 2019
Cited alongside, same era.
Mlperf training benchmark
Peter Mattson, Christine Cheng, Gregory Diamos, Cody Coleman, Paulius Micikevicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, et al · 2020
Later among the works it cites.
Nvidia A100 tensor core GPU architecture, 2020
NVIDIA · 2020
Later among the works it cites.
Do transformers need deep long-range memory?
Jack Rae and Ali Razavi · 2020
Later among the works it cites.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap · 2020
Later among the works it cites.
XLA: Compiling machine learning for peak performance
Amit Sabne · 2020
Later among the works it cites.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander M Rush · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Stabilizing the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M Roy, and Michael Carbin · 2019
Cited alongside, same era.
Openwebtext corpus, 2019
Aaron Gokaslan, Vanya Cohen, Pavlick Ellie, and Stefanie Tellex · 2019
Cited alongside, same era.
Dissecting the graphcore IPU architecture via microbenchmarking
Zhe Jia, Blake Tillman, Marco Maggioni, and Daniele Paolo Scarpazza · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Later among the works it cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Later among the works it cites.
LambdaNetworks: Modeling long-range interactions without attention
Irwan Bello · 2021
Later among the works it cites.
Paragraph-level rationale extraction through regularization: A case study on european court of human rights cases
Ilias Chalkidis, Manos Fergadiotis, Dimitrios Tsarapatsanis, Nikolaos Aletras, Ion Androutsopoulos, and Prodromos Malakasiotis · 2021
Later among the works it cites.
Kernel operations on the gpu, with autodiff, without memory overflows
Benjamin Charlier, Jean Feydy, Joan Alexis Glaunès, François-David Collin, and Ghislain Durif · 2021
Later among the works it cites.
Scatterbrain: Unifying sparse and low-rank attention
Beidi Chen, Tri Dao, Eric Winsor, Zhao Song, Atri Rudra, and Christopher Ré · 2021
Later among the works it cites.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré · 2021
Later among the works it cites.
Data movement is all you need: A case study on optimizing transformers
Andrei Ivanov, Nikoli Dryden, Tal Ben-Nun, Shigang Li, and Torsten Hoefler · 2021
Later among the works it cites.
Dissecting the Ampere GPU architecture via microbenchmarking
Zhe Jia and Peter Van Sandt · 2021
Later among the works it cites.
Luna: Linear unified nested attention
Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma, and Luke Zettlemoyer · 2021
Later among the works it cites.
Self-attention does not need O ( n 2 ) {O}(n^{2}) memory
Markus N Rabe and Charles Staats · 2021
Later among the works it cites.
Combiner: Full attention transformer with sparse computation cost
Hongyu Ren, Hanjun Dai, Zihang Dai, Mengjiao Yang, Jure Leskovec, Dale Schuurmans, and Bo Dai · 2021
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier · 2021
Later among the works it cites.
Nyströmformer: A nystöm-based algorithm for approximating self-attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
Shuangfei Zhai, Walter Talbott, Nitish Srivastava, Chen Huang, Hanlin Goh, Ruixiang Zhang, and Josh Susskind · 2021
Later among the works it cites.
Long-short transformer: Efficient transformers for language and vision
Chen Zhu, Wei Ping, Chaowei Xiao, Mohammad Shoeybi, Tom Goldstein, Anima Anandkumar, and Bryan Catanzaro · 2021
Later among the works it cites.
Revisiting transformer-based models for long document classification
Xiang Dai, Ilias Chalkidis, Sune Darkner, and Desmond Elliott · 2022
Closest in time.
It’s raw! audio generation with state-space models
Karan Goel, Albert Gu, Chris Donahue, and Christopher Ré · 2022
Closest in time.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Ré · 2022
Closest in time.
Transformer quality in linear time
Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc V Le · 2022
Closest in time.
Nvidia H100 tensor core GPU architecture, 2022
NVIDIA · 2022
Closest in time.
Deepnet: Scaling transformers to 1,000 layers
Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Dongdong Zhang, and Furu Wei · 2022
Closest in time.