Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Original
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Later among the works it cites.
Neural controlled differential equations for irregular time series
Original
Patrick Kidger, James Morrill, James Foster, and Terry Lyons · 2020
Later among the works it cites.
Time-aware large kernel convolutions
Vasileios Lioutas and Yuhong Guo · 2020
Later among the works it cites.
Parallelizing legendre memory unit training
Narsimha Chilkuri and Chris Eliasmith · 2021
Closest in time.
Lipschitz recurrent neural networks
N Benjamin Erichson, Omri Azencot, Alejandro Queiruga, Liam Hodgkinson, and Michael W Mahoney · 2021
Closest in time.
Combining recurrent, convolutional, and continuous-time models with the structured learnable linear state space layer
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré · 2021
Closest in time.
Ckconv: Continuous kernel convolution for sequential data
Original
David W Romero, Anna Kuzina, Erik J Bekkers, Jakub M Tomczak, and Mark Hoogendoorn · 2021
Closest in time.
Unicornn: A recurrent model for learning very long time dependencies
T Konstantin Rusch and Siddhartha Mishra · 2021
Closest in time.
Long range arena : A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Closest in time.
Mlp-mixer: An all-mlp architecture for vision
Original
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, et al · 2021
Closest in time.
Informer: Beyond efficient transformer for long sequence time-series forecasting
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang · 2021
Closest in time.
It’s raw! audio generation with state-space models
Original
Karan Goel, Albert Gu, Chris Donahue, and Christopher Ré · 2022
Closest in time.
Flexconv: Continuous kernel convolutions with differentiable kernel sizes
David W Romero, Robert-Jan Bruintjes, Jakub M Tomczak, Erik J Bekkers, Mark Hoogendoorn, and Jan C van Gemert · 2022
Closest in time.