Fetching the paper…
Reading the bibliography…
Transformers are arguably the main workhorse in recent Natural Language Processing research.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Non-autoregressive transformer by position learning
Yu Bao, Hao Zhou, Jiangtao Feng, Mingxuan Wang, Shujian Huang, Jiajun Chen, and Lei Li · 1911
Earlier work this paper cites.
TENER: adapting transformer encoder for named entity recognition
Hang Yan, Bocao Deng, Xiaonan Li, and Xipeng Qiu · 1911
Earlier work this paper cites.
Graph-bert: Only attention is needed for learning graph representations
Jiawei Zhang, Haopeng Zhang, Congying Xia, and Li Sun · 2001
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
A convolutional neural network for modelling sentences
Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Disan: Directional self-attention network for rnn/cnn-free language understanding
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, Shirui Pan, and Chengqi Zhang · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
An augmented transformer architecture for natural language generation tasks
Hailiang Li, Adele Y. C. Wang, Yang Liu, Du Tang, Zhibin Lei, and Wenye Li · 2019
Earlier work this paper cites.
On the relation between position information and sentence length in neural machine translation
Masato Neishi and Naoki Yoshinaga · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Analysis of positional encodings for neural machine translation
Jan Rosendahl, Viet Anh Khoa Tran, Weiyue Wang, and Hermann Ney · 2019
Cited alongside, same era.
Novel positional encodings to enable tree-based transformers
Vighnesh Leonardo Shiv and Chris Quirk · 2019
Cited alongside, same era.
Positional encoding to control output sequence length
Sho Takase and Naoaki Okazaki · 2019
Cited alongside, same era.
Self-attention with structural position representations
Xing Wang, Zhaopeng Tu, Longyue Wang, and Shuming Shi · 2019
Cited alongside, same era.
Assessing the ability of self-attention networks to learn word order
Baosong Yang, Longyue Wang, Derek F. Wong, Lidia S. Chao, and Zhaopeng Tu · 2019
Cited alongside, same era.
Modeling graph structure in transformer for better AMR-to-text generation
What do position embeddings learn? an empirical study of pre-trained language model positional encoding
Yu-An Wang and Yun-Nung Chen · 2020
Later among the works it cites.
Convolutions and self-attention: Re-interpreting relative positions in pre-trained language models
Tyler Chang, Yifan Xu, Weijian Xu, and Zhuowen Tu · 2021
Closest in time.
Demystifying the better performance of position encoding variants for transformer
Pu-Chin Chen, Henry Tsai, Srinadh Bhojanapalli, Hyung Won Chung, Yin-Wen Chang, and Chun-Sung Ferng · 2021
Closest in time.
Rethinking attention with performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J Colwell, and Adrian Weller · 2021
Closest in time.
Distributed representations for multilingual language processing
Philipp Dufter · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jie Zhu, Junhui Li, Muhua Zhu, Longhua Qian, Min Zhang, and Guodong Zhou · 2019
Cited alongside, same era.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Graph transformer for graph-to-sequence learning
Deng Cai and Wai Lam · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Cited alongside, same era.
Self-attention with cross-lingual position representation
Liang Ding, Longyue Wang, and Dacheng Tao · 2020
Cited alongside, same era.
Increasing learning efficiency of self-attention networks through direct position interactions, learnable temperature, and convoluted attention
Philipp Dufter, Martin Schmitt, and Hinrich Schütze · 2020
Cited alongside, same era.
A generalization of transformer networks to graphs
Vijay Prakash Dwivedi and Xavier Bresson · 2021
Closest in time.
DeBERTa: decoding-enhanced BERT with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Closest in time.
Rethinking positional encoding in language pre-training
Guolin Ke, Di He, and Tie-Yan Liu · 2021
Closest in time.
CAPE: encoding relative positions with continuous augmented positional embeddings
Tatiana Likhomanenko, Qiantong Xu, Ronan Collobert, Gabriel Synnaeve, and Alex Rogozhnikov · 2021
Closest in time.
Improving zero-shot translation by disentangling positional information
Danni Liu, Jan Niehues, James Cross, Francisco Guzmán, and Xian Li · 2021
Closest in time.
On the importance of word order information in cross-lingual sequence labeling
Zihan Liu, Genta Indra Winata, Samuel Cahyawijaya, Andrea Madotto, Zhaojiang Lin, and Pascale Fung · 2021
Closest in time.
Relative positional encoding for Transformers with linear complexity
Antoine Liutkus, Ondřej Cífka, Shih-Lun Wu, Umut Şimşekli, Yi-Hsuan Yang, and Gaël Richard · 2021
Closest in time.
Shortformer: Better language modeling using shorter inputs
Ofir Press, Noah A. Smith, and Mike Lewis · 2021
Closest in time.
Modeling graph structure via relative position for text generation from knowledge graphs
Martin Schmitt, Leonardo F. R. Ribeiro, Philipp Dufter, Iryna Gurevych, and Hinrich Schütze · 2021
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Closest in time.
Long range arena : A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Closest in time.
On position embeddings in BERT
Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang, Hao Yang, Qun Liu, and Jakob Grue Simonsen · 2021
Closest in time.
Da-transformer: Distance-aware transformer
Chuhan Wu, Fangzhao Wu, and Yongfeng Huang · 2021
Closest in time.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2074
Closest in time.