Fetching the paper…
Reading the bibliography…
Recently, a new line of works has emerged to understand and improve self-attention in Transformers by treating it as a kernel machine.
Functions of positive and negative type, and their connection with the theory of integral equations
James Mercer · 1909
Earlier work this paper cites.
Linear systems in self-adjoint form
Cornelius Lanczos · 1958
Earlier work this paper cites.
Extensions of lipschitz maps into a Hilbert space
Joram Lindenstrauss and William B. Johnson · 1984
Earlier work this paper cites.
An overview of statistical learning theory
Vladimir N Vapnik · 1999
Earlier work this paper cites.
Least Squares Support Vector Machines
Johan A.K. Suykens, Tony Van Gestel, Joseph De Brabanter, Bart De Moor, and Joos PL Vandewalle · 2002
Earlier work this paper cites.
Linear algebra and its applications
Gilbert Strang · 2006
Earlier work this paper cites.
Reproducing kernel banach spaces for machine learning
Haizhang Zhang, Yuesheng Xu, and Jun Zhang · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Construction of pairs of reproducing kernel banach spaces
Pando G Georgiev, Luis Sánchez-González, and Panos M Pardalos · 2013
Earlier work this paper cites.
The acl anthology network corpus
Dragomir R Radev, Pradeep Muthukrishnan, Vahed Qazvinian, and Amjad Abu-Jbara · 2013
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
SVD revisited: A new variational principle, compatible feature maps and nonlinear extensions
Johan A.K. Suykens · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The uea multivariate time series classification archive, 2018
Anthony Bagnall, Hoang Anh Dau, Jason Lines, Michael Flynn, James Large, Aaron Bostrom, Paul Southam, and Eamonn Keogh · 2018
Earlier work this paper cites.
Listops: A diagnostic dataset for latent tree learning
Nikita Nangia and Samuel Bowman · 2018
Earlier work this paper cites.
Learning long-range spatial dependencies with horizontal gated recurrent units
Drew Linsley, Junkyung Kim, Vijay Veerabadran, Charles Windolf, and Thomas Serre · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Cited alongside, same era.
Transformer dissection: An unified understanding for transformer’s attention via the lens of kernel
Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
XCiT: Cross-covariance image transformers
Alaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski, Matthijs Douze, Armand Joulin, Ivan Laptev, Natalia Neverova, Gabriel Synnaeve, Jakob Verbeek, et al · 2021
Later among the works it cites.
Transformers are deep infinite-dimensional non-Mercer binary kernel machines
Matthew A Wright and Joseph E Gonzalez · 2021
Later among the works it cites.
A transformer-based framework for multivariate time series representation learning
George Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty, and Carsten Eickhoff · 2021
Later among the works it cites.
You only sample (almost) once: Linear cost self-attention via bernoulli sampling
Zhanpeng Zeng, Yunyang Xiong, Sathya Ravi, Shailesh Acharya, Glenn M Fung, and Vikas Singh · 2021
Later among the works it cites.
Soft: Softmax-free transformer with linear complexity
Jiachen Lu, Jinghan Yao, Junge Zhang, Xiatian Zhu, Hang Xu, Weiguo Gao, Chunjing Xu, Tao Xiang, and Li Zhang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Multiscale vision transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Cited alongside, same era.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Later among the works it cites.
Nyströmformer: A nyström-based algorithm for approximating self-attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh · 2021
Later among the works it cites.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong · 2021
Later among the works it cites.
Skyformer: Remodel self-attention with gaussian kernel and nyström method
Yifan Chen, Qi Zeng, Heng Ji, and Yun Yang · 2021
Later among the works it cites.
Flowformer: Linearizing transformers with conservation flows
Haixu Wu, Jialong Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long · 2022
Later among the works it cites.
Improving transformers with probabilistic attention keys
Tam Minh Nguyen, Tan Minh Nguyen, Dung DD Le, Duy Khuong Nguyen, Viet-Anh Tran, Richard Baraniuk, Nhat Ho, and Stanley Osher · 2022
Later among the works it cites.
Fourierformer: Transformer meets generalized fourier integral theorem
Tan Nguyen, Minh Pham, Tam Nguyen, Khai Nguyen, Stanley Osher, and Nhat Ho · 2022
Later among the works it cites.
Kerple: Kernelized relative positional embedding for length extrapolation
Ta-Chung Chi, Ting-Han Fan, Peter J Ramadge, and Alexander Rudnicky · 2022
Later among the works it cites.
On reproducing kernel banach spaces: Generic definitions and unified framework of constructions
Rong Rong Lin, Hai Zhang Zhang, and Jun Zhang · 2022
Later among the works it cites.
cosformer: Rethinking softmax in attention
Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong, and Yiran Zhong · 2022
Later among the works it cites.
Jigsaw-vit: Learning jigsaw puzzles in vision transformer
Yingyi Chen, Xi Shen, Yahui Liu, Qinghua Tao, and Johan A.K. Suykens · 2023
Closest in time.
A primal-dual framework for transformers and neural networks
Tan Minh Nguyen, Tam Minh Nguyen, Nhat Ho, Andrea L. Bertozzi, Richard Baraniuk, and Stanley Osher · 2023
Closest in time.