Fetching the paper…
Reading the bibliography…
Position representation is crucial for building position-aware representations in Transformers.
BLEU: a Method for Automatic Evaluation of Machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Efficient Transformers: A Survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020 · 2009
Earlier work this paper cites.
Deep Learning , chapter 7.4. MIT Press
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Rethinking the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Earlier work this paper cites.
Stronger Baselines for Trustable Results in Neural Machine Translation
Michael Denkowski and Graham Neubig. 2017 · 2017
Earlier work this paper cites.
Convolutional Sequence to Sequence Learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. 2017 · 2017
Earlier work this paper cites.
OpenNMT: Open-Source Toolkit for Neural Machine Translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander Rush. 2017 · 2017
Earlier work this paper cites.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Scaling Neural Machine Translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Earlier work this paper cites.
A Call for Clarity in Reporting BLEU Scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
Self-Attention with Relative Position Representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Cited alongside, same era.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
On the Relation between Position Information and Sentence Length in Neural Machine Translation
Masato Neishi and Naoki Yoshinaga. 2019 · 2019
Cited alongside, same era.
Facebook FAIR’s WMT19 News Translation Task Submission
Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov. 2019 · 2019
Cited alongside, same era.
fairseq: A Fast, Extensible Toolkit for Sequence Modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Analysis of Positional Encodings for Neural Machine Translation
Incorporating Noisy Length Constraints into Transformer with Length-aware Positional Encodings
Yui Oka, Katsuki Chousa, Katsuhito Sudoh, and Satoshi Nakamura. 2020 · 2020
Later among the works it cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Big Bird: Transformers for Longer Sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020 · 2020
Later among the works it cites.
Rethinking Attention with Performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Łukasz Kaiser, et al. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jan Rosendahl, Viet Anh Khoa Tran, Weiyue Wang, and Hermann Ney. 2019 · 2019
Cited alongside, same era.
Positional encoding to control output sequence length
Sho Takase and Naoaki Okazaki. 2019 · 2019
Cited alongside, same era.
Injecting Numerical Reasoning Skills into Language Models
Mor Geva, Ankit Gupta, and Jonathan Berant. 2020 · 2020
Cited alongside, same era.
Improve Transformer Models with Better Relative Position Embeddings
Zhiheng Huang, Davis Liang, Peng Xu, and Bing Xiang. 2020 · 2020
Cited alongside, same era.
Reformer: The Efficient Transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Cited alongside, same era.
The EOS Decision and Length Extrapolation
Benjamin Newman, John Hewitt, Percy Liang, and Christopher D. Manning. 2020 · 2020
Cited alongside, same era.
Philipp Dufter, Martin Schmitt, and Hinrich Schütze. 2021 · 2021
Closest in time.
A Survey on Document-Level Neural Machine Translation: Methods and Evaluation
Sameen Maruf, Fahimeh Saleh, and Gholamreza Haffari. 2021 · 2021
Closest in time.
Do Transformer Modifications Transfer Across Implementations and Applications?
Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Fevry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, et al. 2021 · 2021
Closest in time.
Using Perturbed Length-aware Positional Encoding for Non-autoregressive Neural Machine Translation
Yui Oka, Katsuhito Sudoh, and Satoshi Nakamura. 2021 · 2021
Closest in time.
On Position Embeddings in BERT
Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang, Hao Yang, Qun Liu, and Jakob Grue Simonsen. 2021 · 2021
Closest in time.
DA-Transformer: Distance-aware Transformer
Chuhan Wu, Fangzhao Wu, and Yongfeng Huang. 2021 · 2068
Closest in time.