Fetching the paper…
Reading the bibliography…
Much recent effort has been invested in non-autoregressive neural machine translation, which appears to be an efficient alternative to state-of-the-art autoregressive machine translation on modern GPUs.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann Dauphin, and Michael Auli · 1901
Earlier work this paper cites.
Insertion-based decoding with automatically inferred generation order
Jiatao Gu, Qi Liu, and Kyunghyun Cho · 1902
Earlier work this paper cites.
Insertion transformer: flexible sequence generation via insertion operations
Mitchell Stern, William Chan, Jamie Ryan Kiros, and Jakob Uszkoreit · 1902
Earlier work this paper cites.
Non-autoregressive machine translation with auxiliary regularization
Yiren Wang, Fei Tian, Di He, Tao Qin, ChengXiang Zhai, and Tie-Yan Liu · 1902
Earlier work this paper cites.
Mask-predict: Parallel decoding of conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke S. Zettlemoyer · 1904
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 1904
Earlier work this paper cites.
Jiatao Gu, Changhan Wang, and Jake Zhao · 1905
Earlier work this paper cites.
A generalized framework of sequence generation with application to undirected sequence models, 2019
Elman Mansimov, Alex Wang, and Kyunghyun Cho · 1905
Earlier work this paper cites.
KERMIT: Generative insertion-based modeling for sequences, 2019
William Chan, Nikita Kitaev, Kelvin Guu, Mitchell Stern, and Jakob Uszkoreit · 1906
Earlier work this paper cites.
Learning deep transformer models for machine translation
Qiang Wang, Bei Li, Tong Xiao, Jingbo Zhu, Changliang Li, Derek F. Wong, and Lidia S. Chao · 1906
Earlier work this paper cites.
Raphael Shu, Jason Lee, Hideki Nakayama, and Kyunghyun Cho · 1908
Earlier work this paper cites.
Hint-based training for non-autoregressive machine translation
Zhuohan Li, Zi Lin, Di He, Fei Tian, Tao Qin, Liwei Wang, and Tie-Yan Liu · 1909
Earlier work this paper cites.
FlowSeq: Non-autoregressive conditional sequence generation with generative flow
Xuezhe Ma, Chunting Zhou, Xian Li, Graham Neubig, and Eduard H. Hovy · 1909
Earlier work this paper cites.
An empirical study of generation order for machine translation
William Chan, Mitchell Stern, Jamie Ryan Kiros, and Jakob Uszkoreit · 1910
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer · 1910
Earlier work this paper cites.
Fast structured decoding for sequence models
Zhiqing Sun, Zhuohan Li, Haoqing Wang, Di He, Zi Lin, and Zhihong Deng · 1910
Earlier work this paper cites.
Guiding non-autoregressive neural machine translation decoding with reordering information
Qiu Ran, Yankai Lin, Peng Li, and Jie Zhou · 1911
Earlier work this paper cites.
Minimizing the bag-of-ngrams difference for non-autoregressive neural machine translation
Chenze Shao, Jinchao Zhang, Yang Feng, Fandong Meng, and Jie Zhou · 1911
Earlier work this paper cites.
Understanding knowledge distillation in non-autoregressive machine translation
Chunting Zhou, Graham Neubig, and Jiatao Gu · 1911
Earlier work this paper cites.
Faster transformer decoding: N-gram masked self-attention, 2020
Ciprian Chelba, Mia Chen, Ankur Bapna, and Noam Shazeer · 2001
Earlier work this paper cites.
Semi-autoregressive training improves mask-predict decoding, 2020b
Marjan Ghazvininejad, Omer Levy, and Luke Zettlemoyer · 2001
Earlier work this paper cites.
Non-autoregressive machine translation with disentangled context transformer
Jungo Kasai, James Cross, Marjan Ghazvininejad, and Jiatao Gu · 2001
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2001
Cited alongside, same era.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer · 2001
Cited alongside, same era.
Xiaoya Li, Yuxian Meng, Arianna Yuan, Fei Wu, and Jiwei Li · 2002
Cited alongside, same era.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Cited alongside, same era.
Aligned cross entropy for non-autoregressive machine translation
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann Dauphin · 2017
Later among the works it cites.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2017
Later among the works it cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2017
Later among the works it cites.
Speeding up neural machine translation decoding by shrinking run-time vocabulary
Xing Shi and Kevin Knight · 2017
Later among the works it cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, and Omer Levy · 2004
Cited alongside, same era.
Non-autoregressive machine translation with latent alignments
Chitwan Saharia, William Chan, Saurabh Saxena, and Mohammad Norouzi · 2004
Cited alongside, same era.
ENGINE: Energy-based inference networks for non-autoregressive machine translation
Lifu Tu, Richard Yuanzhe Pang, Sam Wiseman, and Kevin Gimpel · 2005
Cited alongside, same era.
Improving non-autoregressive neural machine translation with monolingual data
Jiawei Zhou and Phillip Keung · 2005
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2006
Cited alongside, same era.
Optimizing parallel reduction in CUDA, 2007
Mark Harris · 2007
Cited alongside, same era.
A simple, fast, and effective reparameterization of ibm model 2
Chris Dyer, Victor Chahuneau, and Noah A. Smith · 2013
Cited alongside, same era.
Recurrent continuous translation models
Nal Kalchbrenner and Phil Blunsom · 2013
Cited alongside, same era.
Later among the works it cites.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor O. K. Li, and Richard Socher · 2018
Later among the works it cites.
Achieving human parity on automatic Chinese to English news translation, 2018
Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mengnan Li, Shujie Liu, Tie-Yan Liu, Renqian Luo, Arul Menezes, Tao Qin, Frank Seide, Xu Tan, Fei Tian, Lijun Wu, Shuangzhi Wu, Yingce Xia, Dongdong Zhang, Zhirui Zhang, and Ming Zhou · 2018
Later among the works it cites.
Fast decoding in sequence models using discrete latent variables
Łukasz Kaiser, Aurko Roy, Ashish Vaswani, Niki Parmar, Samy Bengio, Jakob Uszkoreit, and Noam Shazeer · 2018
Later among the works it cites.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Jason D. Lee, Elman Mansimov, and Kyunghyun Cho · 2018
Later among the works it cites.
End-to-end non-autoregressive neural machine translation with connectionist temporal classification
Jindřich Libovický and Jindřich Helcl · 2018
Later among the works it cites.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu · 2018
Later among the works it cites.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli · 2018
Later among the works it cites.
A call for clarity in reporting BLEU scores
Matt Post · 2018
Later among the works it cites.
OpenNMT system description for WNMT 2018: 800 words/sec on a single-core CPU
Jean Senellart, Dakun Zhang, Bo Wang, Guillaume Klein, Jean-Pierre Ramatchandirin, Josep Crego, and Alexander Rush · 2018
Later among the works it cites.
Blockwise parallel decoding for deep autoregressive models
Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit · 2018
Later among the works it cites.
Accelerating neural transformer via an average attention network
Biao Zhang, Deyi Xiong, and Jinsong Su · 2018
Later among the works it cites.
Recurrent stacking of layers for compact neural machine translation models
Raj Dabre and Atsushi Fujita · 2019
Later among the works it cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz · 2019
Later among the works it cites.
From research to production and back: Ludicrously fast neural machine translation
Young Jin Kim, Marcin Junczys-Dowmunt, Hany Hassan, Alham Fikri Aji, Kenneth Heafield, Roman Grundkiewicz, and Nikolay Bogoychev · 2019
Later among the works it cites.
Jointly masked sequence-to-sequence model for non-autoregressive neural machine translation
Junliang Guo, Linli Xu, and Enhong Chen · 2020
Closest in time.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A. Smith, and Lingpeng Kong · 2021
Closest in time.