2019

Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine Translation

Guo, Junliang, Tan, Xu, Xu, Linli et al.

Understand

Non-autoregressive translation (NAT) models remove the dependence on previous target tokens and generate all target tokens in parallel, resulting in significant inference speedup but at the cost of inferior translation accuracy compared to autoregressive translation (AT) models.

  • Considering that AT models have higher accuracy and are easier to train than NAT models, and both of them share the same model configurations, a natural idea to improve the accuracy of NAT models is to transfer a well-trained AT model to an NAT model through fine-tuning.
  • However, since AT and NAT models differ greatly in training strategy, straightforward fine-tuning does not work well.
  • In this work, we introduce curriculum learning into fine-tuning for NAT.

Built on

Nothing clear enough to list yet.

Similar

Nothing clear enough to list yet.

Then

Nothing clear enough to list yet.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…