Understanding knowledge distillation in non-autoregressive machine translation
Original
Chunting Zhou, Graham Neubig, and Jiatao Gu. 2019 · 1911
Cited alongside, same era.
Learning from noisy labels with distillation
Yuncheng Li, Jianchao Yang, Yale Song, Liangliang Cao, Jiebo Luo, and Li-Jia Li. 2017 · 1936
Cited alongside, same era.
Distilling the knowledge in a neural network
Original
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
Original
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015 · 2015
Cited alongside, same era.
Guided alignment training for topic-aware neural machine translation
Original
Wenhu Chen, Evgeny Matusov, Shahram Khadivi, and Jan-Thorsten Peter. 2016 · 2016
Cited alongside, same era.
Non-autoregressive neural machine translation
Original
Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Parallel wavenet: Fast high-fidelity speech synthesis
Original
Aaron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche, Edward Lockhart, Luis C Cobo, Florian Stimberg, et al. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Large scale distributed neural network training through online distillation
Original
Rohan Anil, Gabriel Pereyra, Alexandre Passos, Robert Ormandi, George E Dahl, and Geoffrey E Hinton. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Non-autoregressive neural machine translation with enhanced decoder input
Junliang Guo, Xu Tan, Di He, Tao Qin, Linli Xu, and Tie-Yan Liu. 2019a
Cited in the paper.
The lj speech dataset, 2017a. url ttps
Keith Ito
Cited in the paper.