Fetching the paper…
Reading the bibliography…
We describe an efficient hierarchical method to compute attention in the Transformer architecture.
Hierarchical attentional hybrid neural networks for document classification
Jader Abreu, Luis Fred, David Macêdo, and C. Zanchettin. 2019 · 1901
Earlier work this paper cites.
Generating long sequences with sparse transformers
R. Child, Scott Gray, A. Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Stand-alone self-attention in vision models
Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jonathon Shlens. 2019 · 1906
Earlier work this paper cites.
Blockwise self-attention for long document understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Scott Yih, Sinong Wang, and Jie Tang. 2019 · 1911
Earlier work this paper cites.
Axial attention in multidimensional transformers
Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans. 2019 · 1912
Earlier work this paper cites.
A fast algorithm for particle simulations
L Greengard and V Rokhlin. 1987 · 1987
Earlier work this paper cites.
Multilevel matrix multiplication and fast solution of integral equations
A. Brandt and A. A. Lubrecht. 1990 · 1990
Earlier work this paper cites.
Fast algorithms for classical physics
L Greengard. 1994 · 1994
Earlier work this paper cites.
Multipole accelerated preconditioned iterative methods for three-dimensional potential integral equations of the first kind
K. Nabors, T. Korsmeyer, and J. White. 1994 · 1994
Earlier work this paper cites.
Matrix Computation
G.H. Golub and C.F. Van Loan. 1996 · 1996
Earlier work this paper cites.
IES3: A fast integral equation solver for efficient 3-dimensional extraction
S. Kapur and D.E. Long. 1997 · 1997
Earlier work this paper cites.
A precorrected-FFT method for electrostatic analysis of complicated 3D structures
Joel R. Phillips and J. K. White. 1997 · 1997
Earlier work this paper cites.
Numerical linear algebra
L.N. Trefethen and D. Bau. 1997 · 1997
Earlier work this paper cites.
A fast hierarchical algorithm for 3-d capacitance extraction
W. Shi, J. Liu, N. Kakani, and T. Yu. 1998 · 1998
Earlier work this paper cites.
A sparse matrix arithmetic based on h-matrices. part I: Introduction to H-matrices
W. Hackbusch. 1999 · 1999
Earlier work this paper cites.
Foundations of Statistical Natural Language Processing
Chris Manning and Hinrich Schütze. 1999 · 1999
Earlier work this paper cites.
A Multigrid Tutorial
W.L. Briggs, V.E. Henson, and S.F. McCormick. 2000 · 2000
Earlier work this paper cites.
A sparse matrix arithmetic based on H-matrices. part II: Application to multi-dimensional problems
W. Hackbusch. 2000 · 2000
Cited alongside, same era.
Multigrid
Ulrich Trottenberg, Cornelius W. Oosterlee, and Anton Schuller. 2000 · 2000
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2001
Cited alongside, same era.
Efficient content-based sparse attention with routing transformers
Aurko Roy, M. Saffar, Ashish Vaswani, and David Grangier. 2020 · 2003
Cited alongside, same era.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Later among the works it cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Music transformer
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Ian Simon, Curtis Hawthorne, Andrew M. Dai, Matthew D. Hoffman, Monica Dinculescu, and Douglas Eck. 2018 · 2018
Later among the works it cites.
Sharp nearby, fuzzy far away: How neural language models use context
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Later among the works it cites.
Document-level neural machine translation with hierarchical attention networks
Lesly Miculicich, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tom B. Brown, Benjamin Pickman Mann, Nick Ryder, Melanie Subbiah, Jean Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, G. Krüger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric J Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Cited alongside, same era.
Synthesizer: Rethinking self-attention in transformer models
Yi Tay, Dara Bahri, Donald Metzler, D. Juan, Zhe Zhao, and Che Zheng. 2020a · 2005
Cited alongside, same era.
Algorithms in FastImp: A fast and wideband impedance extraction program for complicated 3D geometries
Zhenhai Zhu, Ben Song, and J. K. White. 2005 · 2005
Cited alongside, same era.
Fastsies: a fast stochastic integral equation solver for modeling the rough surface effect
Zhenhai Zhu and J. K. White. 2005 · 2005
Cited alongside, same era.
Masked language modeling for proteins via linearly scalable long-context transformers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Jared Davis, Tamás Sarlós, David Belanger, Lucy J. Colwell, and Adrian Weller. 2020 · 2006
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. 2020 · 2006
Cited alongside, same era.
Efficient transformers: A survey
Yi Tay, M. Dehghani, Dara Bahri, and Donald Metzler. 2020d · 2009
Cited alongside, same era.
Later among the works it cites.
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018 · 2018
Later among the works it cites.
Mesh-tensorflow: Deep learning for supercomputers
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, P. Hawkins, H. Lee, Mingsheng Hong, C. Young, Ryan Sepassi, and Blake A. Hechtman. 2018 · 2018
Later among the works it cites.
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018 · 2018
Later among the works it cites.
Adaptive input representations for neural language modeling
Alexei Baevski and M. Auli. 2019 · 2019
Later among the works it cites.
Attention augmented convolutional networks
I. Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Z. Yang, Yiming Yang, J. Carbonell, Quoc V. Le, and R. Salakhutdinov. 2019 · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Hierarchical transformers for multi-document summarization
Yang Liu and Mirella Lapata. 2019 · 2019
Later among the works it cites.
Etc: Encoding long and structured inputs in transformers
Joshua Ainslie, S. Ontañón, C. Alberti, V. Cvicek, Zachary Kenneth Fisher, Philip Pham, Anirudh Ravula, S. Sanghai, Qifan Wang, and L. Yang. 2020 · 2020
Later among the works it cites.
Generative pretraining from pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. 2020 · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontañón, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020 · 2020
Later among the works it cites.