Fetching the paper…
Reading the bibliography…
Relative positional embeddings (RPE) have received considerable attention since RPEs effectively model the relative distance among tokens and enable length extrapolation.
Metric spaces and positive definite functions
Isaac J Schoenberg · 1938
Earlier work this paper cites.
Harmonic analysis on semigroups: theory of positive definite and related functions , volume 100
Christian Berg, Jens Peter Reus Christensen, and Paul Ressel · 1984
Earlier work this paper cites.
Kernel pca and de-noising in feature spaces
Sebastian Mika, Bernhard Schölkopf, Alex Smola, Klaus-Robert Müller, Matthias Scholz, and Gunnar Rätsch · 1998
Earlier work this paper cites.
The kernel trick for distances
Bernhard Schölkopf · 2000
Earlier work this paper cites.
The spectrum kernel: A string kernel for svm protein classification
Christina Leslie, Eleazar Eskin, and William Stafford Noble · 2001
Earlier work this paper cites.
Kernel k-means: spectral clustering and normalized cuts
Inderjit S Dhillon, Yuqiang Guan, and Brian Kulis · 2004
Earlier work this paper cites.
Kernel regression for image processing and reconstruction
Hiroyuki Takeda, Sina Farsiu, and Peyman Milanfar · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
Music transformer
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M. Dai, Matthew D. Hoffman, Monica Dinculescu, and Douglas Eck · 2019
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Cited alongside, same era.
Transformer dissection: An unified understanding for transformer’s attention via the lens of kernel
Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
A study of bfloat16 for deep learning training
Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi, Dipankar Das, Kunal Banerjee, Sasikanth Avancha, Dharma Teja Vooturi, Nataraj Jammalamadaka, Jianyu Huang, Hector Yuen, et al · 2019
Cape: Encoding relative positions with continuous augmented positional embeddings
Tatiana Likhomanenko, Qiantong Xu, Gabriel Synnaeve, Ronan Collobert, and Alex Rogozhnikov · 2021
Later among the works it cites.
A simple and effective positional encoding for transformers
Pu-Chin Chen, Henry Tsai, Srinadh Bhojanapalli, Hyung Won Chung, Yin-Wen Chang, and Chun-Sung Ferng · 2021
Later among the works it cites.
The case for translation-invariant self-attention in transformer-based language models
Ulme Wennberg and Gustav Eje Henter · 2021
Later among the works it cites.
{DEBERTA}: {DECODING}-{enhanced} {bert} {with} {disentangled} {attention}
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Later among the works it cites.
Rethinking positional encoding in language pre-training
Guolin Ke, Di He, and Tie-Yan Liu · 2021
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning · 2020
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Cited alongside, same era.
Improve transformer models with better relative position embeddings
Zhiheng Huang, Davis Liang, Peng Xu, and Bing Xiang · 2020
Cited alongside, same era.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush · 2020
Cited alongside, same era.
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Later among the works it cites.
Rethinking and improving relative position encoding for vision transformer
Kan Wu, Houwen Peng, Minghao Chen, Jianlong Fu, and Hongyang Chao · 2021
Later among the works it cites.
Stable, fast and accurate: Kernelized attention with relative positional encoding
Shengjie Luo, Shanda Li, Tianle Cai, Di He, Dinglan Peng, Shuxin Zheng, Guolin Ke, Liwei Wang, and Tie-Yan Liu · 2021
Later among the works it cites.
Nyströmformer: A nystöm-based algorithm for approximating self-attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh · 2021
Later among the works it cites.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong · 2021
Later among the works it cites.
Rethinking attention with performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J Colwell, and Adrian Weller · 2021
Later among the works it cites.
GPT-NeoX: Large scale autoregressive language modeling in pytorch, 2021
Sid Black, Stella Biderman, Alex Andonian, Quentin Anthony, Preetham Gali, Leo Gao, Eric Hallahan, Josh Levy-Kramer, Connor Leahy, Lucas Nestler, Kip Parker, Jason Phang, Michael Pieler, Shivanshu Purohit, Tri Songz, Phil Wang, and Samuel Weinbach · 2021
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah Smith, and Mike Lewis · 2022
Closest in time.
Position Information in Transformers: An Overview
Philipp Dufter, Martin Schmitt, and Hinrich Schütze · 2022
Closest in time.
cosformer: Rethinking softmax in attention
Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong, and Yiran Zhong · 2022
Closest in time.
DA-transformer: Distance-aware transformer
Chuhan Wu, Fangzhao Wu, and Yongfeng Huang · 2068
Closest in time.