Fetching the paper…
Reading the bibliography…
Neural Machine Translation (NMT) models achieve state-of-the-art performance on many translation benchmarks.
Multilingual neural machine translation with knowledge distillation
Xu Tan, Yi Ren, Di He, Tao Qin, Zhou Zhao, and Tie-Yan Liu. 2019 · 1902
Earlier work this paper cites.
Competence-based curriculum learning for neural machine translation
Emmanouil Antonios Platanios, Otilia Stretcu, Graham Neubig, Barnabas Poczos, and Tom M Mitchell. 2019 · 1903
Earlier work this paper cites.
Distilling task-specific knowledge from bert into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019 · 1903
Earlier work this paper cites.
Dynamic past and future for neural machine translation
Zaixiang Zheng, Shujian Huang, Zhaopeng Tu, Xin-Yu Dai, and Jiajun Chen. 2019 · 1904
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 1908
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Bin Dong, Jikai Hou, Yiping Lu, and Zhihua Zhang. 2019 · 1910
Earlier work this paper cites.
Explaining sequence-level knowledge distillation as data-augmentation for neural machine translation
Mitchell A Gordon and Kevin Duh. 2019 · 1912
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Synthesizer: Rethinking self-attention in transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. 2020 · 2005
Earlier work this paper cites.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen John Maybank, and Dacheng Tao. 2020 · 2006
Earlier work this paper cites.
Norm-based curriculum learning for neural machine translation
Xuebo Liu, Houtim Lai, Derek F Wong, and Lidia S Chao. 2020 · 2006
Cited alongside, same era.
Ternarybert: Distillation-aware ultra-low bit bert
Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu. 2020 · 2009
Cited alongside, same era.
Why skip if you can combine: A simple knowledge distillation technique for intermediate layers
Yimeng Wu, Peyman Passban, Mehdi Rezagholizade, and Qun Liu. 2020 · 2010
Cited alongside, same era.
Learning light-weight translation models from deep transformer
Bei Li, Ziyang Wang, Hui Liu, Quan Du, Tong Xiao, Chunliang Zhang, and Jingbo Zhu. 2020 · 2012
Cited alongside, same era.
Sequence to sequence learning with neural networks
Curriculum learning and minibatch bucketing in neural machine translation
Tom Kocmi and Ondrej Bojar. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Attention-guided answer distillation for machine reading comprehension
Minghao Hu, Yuxing Peng, Furu Wei, Zhen Huang, Dongsheng Li, Nan Yang, and Ming Zhou. 2018 · 2018
Later among the works it cites.
DTMT: A novel deep transition architecture for neural machine translation
Fandong Meng and Jinchao Zhang. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Cited alongside, same era.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Ensemble distillation for neural machine translation
Markus Freitag, Yaser Al-Onaizan, and Baskaran Sankaran. 2017 · 2017
Cited alongside, same era.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Distilling knowledge learned in bert for text generation
Yen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu, and Jingjing Liu. 2020b
Cited in the paper.
Future-aware knowledge distillation for neural machine translation
Biao Zhang, Deyi Xiong, Jinsong Su, and Jiebo Luo. 2019a
Cited in the paper.
The evolved transformer
David So, Quoc Le, and Chen Liang. 2019 · 2019
Later among the works it cites.
Online distilling from checkpoints for neural machine translation
Hao-Ran Wei, Shujian Huang, R. Wang, Xin-Yu Dai, and Jiajun Chen. 2019 · 2019
Later among the works it cites.
Bridging the gap between prior and posterior knowledge selection for knowledge-grounded dialogue generation
Xiuyi Chen, Fandong Meng, Peng Li, Feilong Chen, Shuang Xu, Bo Xu, and Jie Zhou. 2020a · 2020
Later among the works it cites.
WeChat neural machine translation systems for WMT20
Fandong Meng, Jianhao Yan, Yijin Liu, Yuan Gao, Xianfeng Zeng, Qinsong Zeng, Peng Li, Ming Chen, Jie Zhou, Sifan Liu, and Hao Zhou. 2020 · 2020
Later among the works it cites.
Multi-unit transformers for neural machine translation
Jianhao Yan, Fandong Meng, and Jie Zhou. 2020 · 2020
Later among the works it cites.