Fetching the paper…
Reading the bibliography…
How do we perform efficient inference while retaining high translation quality? Existing neural machine translation models, such as Transformer, achieve high performance, but they decode words one by one, which is inefficient.
Fitting autoregressive models for prediction
Akaike, H. 1969 · 1969
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
Glancing Transformer for non-autoregressive neural machine translation
Qian, L.; Zhou, H.; Bao, Y.; Wang, M.; Qiu, L.; Zhang, W.; Yu, Y.; and Li, L. 2021 · 2003
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A.; Fernández, S.; Gomez, F.; and Schmidhuber, J. 2006 · 2006
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D.; Cho, K.; and Bengio, Y. 2015 · 2015
Earlier work this paper cites.
Attention is all you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Non-autoregressive neural machine translation
Gu, J.; Bradbury, J.; Xiong, C.; Li, V. O.; and Socher, R. 2018 · 2018
Earlier work this paper cites.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Lee, J.; Mansimov, E.; and Cho, K. 2018 · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M.; Shazeer, N.; and Uszkoreit, J. 2018 · 2018
Earlier work this paper cites.
Semi-autoregressive neural machine translation
Wang, C.; Zhang, J.; and Chen, H. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Mask-predict: parallel decoding of conditional masked language models
Ghazvininejad, M.; Levy, O.; Liu, Y.; and Zettlemoyer, L. 2019 · 2019
Cited alongside, same era.
Levenshtein Transformer
Gu, J.; Wang, C.; and Zhao, J. 2019 · 2019
Cited alongside, same era.
FairSeq: a fast, extensible toolkit for sequence modeling
Ott, M.; Edunov, S.; Baevski, A.; Fan, A.; Gross, S.; Ng, N.; Grangier, D.; and Auli, M. 2019 · 2019
Cited alongside, same era.
Insertion Transformer: flexible sequence generation via insertion operations
Stern, M.; Chan, W.; Kiros, J.; and Uszkoreit, J. 2019 · 2019
Depth-adaptive Transformer
Elbayad, M.; Gu, J.; Grave, E.; and Auli, M. 2020 · 2020
Later among the works it cites.
Aligned cross entropy for non-autoregressive machine translation
Ghazvininejad, M.; Karpukhin, V.; Zettlemoyer, L.; and Levy, O. 2020 · 2020
Later among the works it cites.
Improving Transformer optimization through better initialization
Huang, X. S.; Perez, F.; Ba, J.; and Volkovs, M. 2020 · 2020
Later among the works it cites.
Non-autoregressive machine translation with latent alignments
Saharia, C.; Chan, W.; Saxena, S.; and Norouzi, M. 2020 · 2020
Later among the works it cites.
Understanding knowledge distillation in non-autoregressive machine translation
Zhou, C.; Gu, J.; and Neubig, G. 2020 · 2020
Later among the works it cites.
Non-autoregressive translation by learning target categorical codes
Bao, Y.; Huang, S.; Xiao, T.; Wang, D.; Dai, X.; and Chen, J. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fast structured decoding for sequence models
Sun, Z.; Li, Z.; Wang, H.; He, D.; Lin, Z.; and Deng, Z. 2019 · 2019
Cited alongside, same era.
Imitation learning for non-autoregressive neural machine translation
Wei, B.; Wang, M.; Zhou, H.; Lin, J.; and Sun, X. 2019 · 2019
Cited alongside, same era.
Imputer: Sequence modelling via imputation and dynamic programming
Chan, W.; Saharia, C.; Hinton, G.; Norouzi, M.; and Jaitly, N. 2020 · 2020
Cited alongside, same era.
Closest in time.
Fully non-autoregressive neural machine translation: tricks of the trade
Gu, J.; and Kong, X. 2021 · 2021
Closest in time.
Deep encoder, shallow decoder: reevaluating non-autoregressive machine translation
Kasai, J.; Pappas, N.; Peng, H.; Cross, J.; and Smith, N. 2021 · 2021
Closest in time.