Fetching the paper…
Reading the bibliography…
Dot-product attention mechanism plays a crucial role in modern deep architectures (e.g., Transformer) for sequence modeling, however, na\"ive exact computation of this model incurs quadratic time and memory complexities in sequence length, hindering the training of long-sequence models.
Maintaining discrete probability distributions optimally
Torben Hagerup, Kurt Mehlhorn, and J Ian Munro · 1993
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Bernhard Schölkopf, Alexander J Smola, Francis Bach, et al · 2002
Earlier work this paper cites.
Comparing distributions and shapes using the kernel distance
Sarang Joshi, Raj Varma Kommaraji, Jeff M Phillips, and Suresh Venkatasubramanian · 2011
Earlier work this paper cites.
Randomized primitives for linear algebra and applications
Anastasios Zouzias · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
An introduction to matrix concentration inequalities
Joel A Tropp · 2015
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
Hashing-based-estimators for kernel density in high dimensions
Moses Charikar and Paris Siminelakis · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Efficient density evaluation for smooth kernels
Arturs Backurs, Moses Charikar, Piotr Indyk, and Paris Siminelakis · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Space and time efficient kernel density estimation in high dimensions
Arturs Backurs, Piotr Indyk, and Tal Wagner · 2019
Cited alongside, same era.
Large Scale GAN Training for High Fidelity Natural Image Synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2019
Cited alongside, same era.
Rehashing kernel evaluation in high dimensions
Paris Siminelakis, Kexin Rong, Peter Bailis, Moses Charikar, and Philip Levis · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Later among the works it cites.
Rethinking Attention with Performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2021
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Kernel density estimation through density constrained near neighbor search
Moses Charikar, Michael Kapralov, Navid Nouri, and Paris Siminelakis · 2020
Cited alongside, same era.
MONGOOSE: A learnable LSH framework for efficient neural network training
Beidi Chen, Zichang Liu, Binghui Peng, Zhaozhuo Xu, Jonathan Lingjie Li, Tri Dao, Zhao Song, Anshumali Shrivastava, and Christopher Re · 2020
Cited alongside, same era.
Smyrf-efficient attention using asymmetric clustering
Giannis Daras, Nikita Kitaev, Augustus Odena, and Alexandros G Dimakis · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Francois Fleuret · 2020
Cited alongside, same era.
Reformer: The Efficient Transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
Beidi Chen, Tri Dao, Kaizhao Liang, Jiaming Yang, Zhao Song, Atri Rudra, and Christopher Re
Cited in the paper.
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier · 2021
Later among the works it cites.
Sparse Attention with Learning to Hash
Zhiqing Sun, Yiming Yang, and Shinjae Yoo · 2021
Later among the works it cites.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
Informer: Beyond efficient transformer for long sequence time-series forecasting
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang · 2021
Later among the works it cites.
Jyotikrishna Dass, Shang Wu, Huihong Shi, Chaojian Li, Zhifan Ye, Zhongfeng Wang, and Yingyan Lin · 2022
Later among the works it cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2022
Later among the works it cites.