Fetching the paper…
Reading the bibliography…
Conditional masked language model (CMLM) training has proven successful for non-autoregressive and semi-autoregressive sequence generation tasks, such as machine translation.
A generalized framework of sequence generation with application to undirected sequence models
Elman Mansimov, Alex Wang, and Kyunghyun Cho. 2019 · 1905
Earlier work this paper cites.
Semi-autoregressive training improves mask-predict decoding
Marjan Ghazvininejad, Omer Levy, and Luke Zettlemoyer. 2020 · 2001
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2010
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Ondrej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Ale s Tamchyna. 2014 · 2014
Earlier work this paper cites.
Findings of the 2017 conference on machine translation (wmt17)
Ond rej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, and Marco Turchi. 2017 · 2017
Earlier work this paper cites.
Learning what’s easy: Fully differentiable neural easy-first taggers
André F. T. Martins and Julia Kreutzer. 2017 · 2017
Cited alongside, same era.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor O.K. Li, and Richard Socher. 2018 · 2018
Cited alongside, same era.
Fast decoding in sequence models using discrete latent variables
Lukasz Kaiser, Samy Bengio, Aurko Roy, Ashish Vaswani, Niki Parmar, Jakob Uszkoreit, and Noam Shazeer. 2018 · 2018
Cited alongside, same era.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Jason Lee, Elman Mansimov, and Kyunghyun Cho. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Mask-predict: Parallel decoding of conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer. 2019 · 2019
Later among the works it cites.
Levenshtein transformer
Jiatao Gu, Changhan Wang, and Junbo Zhao. 2019 · 2019
Later among the works it cites.
Insertion transformer: Flexible sequence generation via insertion operations
Mitchell Stern, Will Chan, Jamie Kiros, and Jakob Uszkoreit. 2019 · 2019
Later among the works it cites.
Raphael Shu, Jason Lee, Hideki Nakayama, and Kyunghyun Cho. 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…