Fetching the paper…
Reading the bibliography…
We present the Insertion Transformer, an iterative, partially autoregressive model for sequence generation based on insertion operations.
Breaking the Softmax Bottleneck: A High-Rank RNN Language Model
Yang, Z., Dai, Z., Salakhutdinov, R., and Cohen, W. W · 1902
Earlier work this paper cites.
Sequence Transduction with Recurrent Neural Networks
Graves, A · 2012
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks
Sutskever, I., Vinyals, O., and Le, Q · 2014
Earlier work this paper cites.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation
Luong, M.-T., Pham, H., and Manning, C. D · 2015
Earlier work this paper cites.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhutdinov, R., Zemel, R., and Bengio, Y · 2015
Earlier work this paper cites.
End-to-End Attention-based Large Vocabulary Speech Recognition
Bahdanau, D., Chorowski, J., Serdyuk, D., Brakel, P., and Bengio, Y · 2016
Cited alongside, same era.
Listen, Attend and Spell: A Neural Network for Large Vocabulary Conversational Speech Recognition
Chan, W., Jaitly, N., Le, Q., and Vinyals, O · 2016
Cited alongside, same era.
Sequence-Level Knowledge Distillation
Kim, Y. and Rush, A. M · 2016
Cited alongside, same era.
Reward Augmented Maximum Likelihood for Neural Structured Prediction
Norouzi, M., Bengio, S., Zhifeng Chen, N. J., Schuster, M., Wu, Y., and Schuurmans, D · 2016
Cited alongside, same era.
Policy Distillation
Rusu, A. A., Colmenarejo, S. G., Gulcehre, C., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., and Hadsell, R · 2016
Cited alongside, same era.
Parallel Multiscale Autoregressive Density Estimation
Reed, S., van den Oord, A., Kalchbrenner, N., Colmenarejo, S. G., Wang, Z., Belov, D., and de Freitas, N · 2017
Non-Autoregressive Neural Machine Translation
Gu, J., Bradbury, J., Xiong, C., Li, V. O., and Socher, R · 2018
Later among the works it cites.
Deterministic Non-Autoregressive Neural Sequence Modeling by Iterative Refinement
Lee, J., Mansimov, E., and Cho, K · 2018
Later among the works it cites.
Generating Sentences Using a Dynamic Canvas
Shah, H., Zheng, B., and Barber, D · 2018
Later among the works it cites.
Blockwise Parallel Decoding for Deep Autoregressive Models
Stern, M., Shazeer, N., and Uszkoreit, J · 2018
Later among the works it cites.
Tensor2Tensor for Neural Machine Translation
Vaswani, A., Bengio, S., Brevdo, E., Chollet, F., Gomez, A. N., Gouws, S., Jones, L., Kaiser, L., Kalchbrenner, N., Parmar, N., Sepassi, R., Shazeer, N., and Uszkoreit, J · 2018
Later among the works it cites.
Semi-Autoregressive Neural Machine Translation
Wang, C., Zhang, J., and Chen, H · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attention Is All You Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Tacotron: Towards End-to-End Speech Synthesis
Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R. J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., Le, Q., Agiomyrgiannakis, Y., Clark, R., and Saurous, R. A · 2017
Cited alongside, same era.
The Importance of Generation Order in Language Modeling
Ford, N., Duckworth, D., Norouzi, M., and Dahl, G. E · 2018
Cited alongside, same era.
WaveNet: A Generative Model for Raw Audio
Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K
Cited in the paper.
Pixel Recurrent Neural Networks
Oord, A., Kalchbrenner, N., and Kavukcuoglu, K
Cited in the paper.
Conditional Image Generation with PixelCNN Decoders
Oord, A., Kalchbrenner, N., Vinyals, O., Espeholt, L., Graves, A., and Kavukcuoglu, K
Cited in the paper.
Later among the works it cites.
Insertion-based Decoding with Automatically Inferred Generation Order
Gu, J., Liu, Q., and Cho, K · 2019
Closest in time.
Non-Monotonic Sequential Text Generation
Welleck, S., Brantley, K., Daume, H., and Cho, K · 2019
Closest in time.