2019

Depth-Adaptive Transformer

Elbayad, Maha, Gu, Jiatao, Grave, Edouard et al.

Understand

State of the art sequence-to-sequence models for large scale tasks perform a fixed number of computations for each input sequence regardless of whether it is easy or hard to process.

  • In this paper, we train Transformer models which can make output predictions at different stages of the network and we investigate different ways to predict how much computation is required for a particular sequence.
  • Unlike dynamic computation in Universal Transformers, which applies the same set of layers iteratively, we apply different layers at every step to adjust both the amount of computation as well as the model capacity.
  • On IWSLT German-English translation our approach matches the accuracy of a well tuned baseline Transformer while using less than a quarter of the decoder layers.

Built on

  • BLEU: a method for automatic evaluation of machine translation

    K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002

    Earlier work this paper cites.

  • Report on the 11th iwslt evaluation campaign

    M. Cettolo, J. Niehues, S. Stüker, L. Bentivogli, and M. Federico · 2014

    Earlier work this paper cites.

  • Adam: A method for stochastic optimization

    D. Kingma and J. Ba · 2015

    Earlier work this paper cites.

  • Adaptive computation time for recurrent neural networks

    Alex Graves · 2016

    Earlier work this paper cites.

  • Neural machine translation of rare words with subword units

    R. Sennrich, B. Haddow, and A. Birch · 2016

    Earlier work this paper cites.

  • Branchynet: Fast inference via early exiting from deep neural networks

    Surat Teerapittayanon, Bradley McDanel, and Hsiang-Tsung Kung · 2016

    Earlier work this paper cites.

Similar

  • Adaptive neural networks for efficient inference

    Tolga Bolukbasi, Joseph Wang, Ofer Dekel, and Venkatesh Saligrama · 2017

    Cited alongside, same era.

  • Probabilistic adaptive computation time

    Michael Figurnov, Artem Sobolev, and Dmitry P. Vetrov · 2017

    Cited alongside, same era.

  • Convolutional sequence to sequence learning

    Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin · 2017

    Cited alongside, same era.

  • Multi-scale dense networks for resource efficient image classification

    Gao Huang, Danlu Chen, Tianhong Li, Felix Wu, Laurens van der Maaten, and Kilian Q Weinberger · 2017

    Cited alongside, same era.

  • Attention is all you need

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I. Polosukhin · 2017

    Cited alongside, same era.

  • Universal transformers

    Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2018

    Cited alongside, same era.

Then

  • Classical structured prediction losses for sequence to sequence learning

    Sergey Edunov, Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato · 2018

    Later among the works it cites.

  • Skipnet: Learning dynamic routing in convolutional networks

    Xin Wang, Fisher Yu, Zi-Yi Dou, Trevor Darrell, and Joseph E Gonzalez · 2018

    Later among the works it cites.

  • BERT: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019

    Closest in time.

  • Facebook fair’s wmt19 news translation task submission

    Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov · 2019

    Closest in time.

  • Fairseq: A fast, extensible toolkit for sequence modeling

    Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019

    Closest in time.

  • Language models are unsupervised multitask learners

    Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…