Fetching the paper…
Reading the bibliography…
Neural machine translation models are often biased toward the limited translation references seen during training.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Andre Martins and Ramon Astudillo. 2016 · 2016
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Sequence-to-sequence learning as beam-search optimization
Sam Wiseman and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2017 · 2017
Earlier work this paper cites.
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles. 2017 · 2017
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Christopher Maddison, Andriy Mnih, and Yee Teh. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Training neural machine translation using word embedding-based loss
Katsuki Chousa, Katsuhito Sudoh, and Satoshi Nakamura. 2018 · 2018
Earlier work this paper cites.
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018 · 2018
Cited alongside, same era.
Classical structured prediction losses for sequence to sequence learning
Sergey Edunov, Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato. 2018 · 2018
Cited alongside, same era.
Bag-of-words as target for neural machine translation
Shuming Ma, Xu Sun, Yizhong Wang, and Junyang Lin. 2018 · 2018
Cited alongside, same era.
Document-level neural machine translation with hierarchical attention networks
Lesly Miculicich, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Deep reinforcement learning with distributional semantic rewards for abstractive summarization
Siyao Li, Deren Lei, Pengda Qin, and William Yang Wang. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Tied transformers: Neural machine translation with shared encoder and decoder
Yingce Xia, Tianyu He, Xu Tan, Fei Tian, Di He, and Tao Qin. 2019 · 2019
Later among the works it cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Later among the works it cites.
Language model prior for low-resource neural machine translation
Christos Baziotis, Barry Haddow, and Alexandra Birch. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
On the use of bert for neural machine translation
Stéphane Clinchant, Kweon Woo Jung, and Vassilina Nikoulina. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Pre-trained language model representations for language generation
Sergey Edunov, Alexei Baevski, and Michael Auli. 2019 · 2019
Cited alongside, same era.
ReWE: Regressing word embeddings for regularization of neural machine translation systems
Inigo Jauregi Unanue, Ehsan Zare Borzeshi, Nazanin Esmaili, and Massimo Piccardi. 2019 · 2019
Cited alongside, same era.
Von mises-fisher loss for training sequence to sequence models with continuous outputs
Sachin Kumar and Yulia Tsvetkov. 2019 · 2019
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Julien Chaumond, Lysandre Debut, Victor Sanh, Clement Delangue, Anthony Moi, Pierric Cistac, Morgan Funtowicz, Joe Davison, Sam Shleifer, et al. 2020 · 2020
Later among the works it cites.
Sequence generation with mixed representations
Lijun Wu, Shufang Xie, Yingce Xia, Yang Fan, Jian-Huang Lai, Tao Qin, and Tie-Yan Liu. 2020 · 2020
Later among the works it cites.
Towards making the most of bert in neural machine translation
Jiacheng Yang, Mingxuan Wang, Hao Zhou, Chengqi Zhao, Weinan Zhang, Yong Yu, and Lei Li. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Incorporating bert into neural machine translation
Jinhua Zhu, Yingce Xia, Lijun Wu, Di He, Tao Qin, Wengang Zhou, Houqiang Li, and Tie-Yan Liu. 2020 · 2020
Later among the works it cites.
Token-level and sequence-level loss smoothing for RNN language models
Maha Elbayad, Laurent Besacier, and Jakob Verbeek. 2018 · 2094
Closest in time.