Fetching the paper…
Reading the bibliography…
Most sequence-to-sequence (seq2seq) models are autoregressive; they generate each token by conditioning on previously generated tokens.
Insertion-based decoding with automatically inferred generation order
Jiatao Gu, Qi Liu, and Kyunghyun Cho. 2019 · 1902
Earlier work this paper cites.
Macow: Masked convolutional generative flow
Xuezhe Ma and Eduard Hovy. 2019 · 1902
Earlier work this paper cites.
Insertion transformer: Flexible sequence generation via insertion operations
Mitchell Stern, William Chan, Jamie Kiros, and Jakob Uszkoreit. 2019 · 1902
Earlier work this paper cites.
Non-autoregressive machine translation with auxiliary regularization
Yiren Wang, Fei Tian, Di He, Tao Qin, ChengXiang Zhai, and Tie-Yan Liu. 2019 · 1902
Earlier work this paper cites.
Constant-time machine translation with conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer. 2019 · 1904
Earlier work this paper cites.
Raphael Shu, Jason Lee, Hideki Nakayama, and Kyunghyun Cho. 2019 · 1908
Earlier work this paper cites.
Graphical models, exponential families, and variational inference
Martin J Wainwright, Michael I Jordan, et al. 2008 · 2008
Earlier work this paper cites.
Wit3: Web inventory of transcribed and translated talks
Mauro Cettolo, Christian Girardi, and Marcello Federico. 2012 · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Generating sentences from a continuous space
Samuel R Bowman, Luke Vilnis, Oriol Vinyals, Andrew M Dai, Rafal Jozefowicz, and Samy Bengio. 2015 · 2015
Earlier work this paper cites.
Importance weighted autoencoders
Yuri Burda, Roger Grosse, and Ruslan Salakhutdinov. 2015 · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Cited alongside, same era.
Variational inference with normalizing flows
Danilo Jimenez Rezende and Shakir Mohamed. 2015 · 2015
Cited alongside, same era.
A neural attention model for abstractive sentence summarization
Alexander M Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Cited alongside, same era.
Density estimation using real nvp
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. 2016 · 2016
Cited alongside, same era.
Improving variational inference with inverse autoregressive flow
Glow: Generative flow with invertible 1x1 convolutions
Durk P Kingma and Prafulla Dhariwal. 2018 · 2018
Later among the works it cites.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Jason Lee, Elman Mansimov, and Kyunghyun Cho. 2018 · 2018
Later among the works it cites.
End-to-end non-autoregressive neural machine translation with connectionist temporal classification
Jindřich Libovickỳ and Jindřich Helcl. 2018 · 2018
Later among the works it cites.
Stack-pointer networks for dependency parsing
Xuezhe Ma, Zecong Hu, Jingzhou Liu, Nanyun Peng, Graham Neubig, and Eduard Hovy. 2018 · 2018
Later among the works it cites.
Analyzing uncertainty in neural machine translation
Myle Ott, Michael Auli, David Grangier, et al. 2018 · 2018
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik P Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling. 2016 · 2016
Cited alongside, same era.
context2vec: Learning generic context embedding with bidirectional LSTM
Oren Melamud, Jacob Goldberger, and Ido Dagan. 2016 · 2016
Cited alongside, same era.
Masked autoregressive flow for density estimation
George Papamakarios, Theo Pavlakou, and Iain Murray. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher. 2018 · 2018
Cited alongside, same era.
Sequence to sequence mixture model for diverse machine translation
Xuanli He, Gholamreza Haffari, and Mohammad Norouzi. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Mae: Mutual posterior-divergence regularization for variational autoencoders
Xuezhe Ma, Chunting Zhou, and Eduard Hovy. 2019 · 2019
Closest in time.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Closest in time.
Waveglow: A flow-based generative network for speech synthesis
Ryan Prenger, Rafael Valle, and Bryan Catanzaro. 2019 · 2019
Closest in time.
Mixture models for diverse machine translation: Tricks of the trade
Tianxiao Shen, Myle Ott, Michael Auli, et al. 2019 · 2019
Closest in time.
Latent normalizing flows for discrete sequences
Zachary Ziegler and Alexander Rush. 2019 · 2019
Closest in time.