Fetching the paper…
Reading the bibliography…
We study the power of cross-attention in the Transformer architecture within the context of transfer learning for machine translation, and extend the findings of studies into cross-attention when training from scratch.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
A universal parent model for low-resource neural machine translation transfer
Mozhdeh Gheini and Jonathan May. 2019 · 1909
Earlier work this paper cites.
Not all parameters are born equal: Attention is mostly what you need
Nikolay Bogoychev. 2020 · 2010
Earlier work this paper cites.
Takashi Wada, Tomoharu Iwata, Yuji Matsumoto, Timothy Baldwin, and Jey Han Lau. 2020 · 2010
Earlier work this paper cites.
Exploiting similarities among languages for machine translation
Tomás Mikolov, Quoc V. Le, and Ilya Sutskever. 2013 · 2013
Earlier work this paper cites.
An empirical investigation of catastrophic forgeting in gradientbased neural networks
Ian J. Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Normalized word embedding and orthogonal transform for bilingual word translation
Chao Xing, Dong Wang, Chao Liu, and Yiye Lin. 2015 · 2015
Earlier work this paper cites.
Learning principled bilingual mappings of word embeddings while preserving monolingual invariance
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Transfer learning for low-resource neural machine translation
Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Learning bilingual word embeddings with (almost) no bilingual data
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2017 · 2017
Earlier work this paper cites.
Transfer learning across low-resource, related languages for neural machine translation
Toan Q. Nguyen and David Chiang. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Trivial transfer learning for low-resource neural machine translation
Tom Kocmi and Ondřej Bojar. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Word translation without parallel data
Guillaume Lample, Alexis Conneau, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018 · 2018
Cited alongside, same era.
Rapid adaptation of neural machine translation to new languages
Graham Neubig and Junjie Hu. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
When and why are pre-trained word embeddings useful for neural machine translation?
Ye Qi, Devendra Sachan, Matthieu Felix, Sarguna Padmanabhan, and Graham Neubig. 2018 · 2018
Cited alongside, same era.
Simple, scalable adaptation for neural machine translation
Ankur Bapna and Orhan Firat. 2019 · 2019
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Later among the works it cites.
Finding the optimal vocabulary size for neural machine translation
Thamme Gowda and Jonathan May. 2020 · 2020
Later among the works it cites.
Investigating catastrophic forgetting during continual training for neural machine translation
Shuhao Gu and Yang Feng. 2020 · 2020
Later among the works it cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Understanding neural machine translation by simplification: The case of encoder-free models
Gongbo Tang, Rico Sennrich, and Joakim Nivre. 2019 · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
Monolingual adapters for zero-shot neural machine translation
Jerin Philip, Alexandre Berard, Matthias Gallé, and Laurent Besacier. 2020 · 2020
Later among the works it cites.
Hard-coded Gaussian attention for neural machine translation
Weiqiu You, Simeng Sun, and Mohit Iyyer. 2020 · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Closest in time.
WARP: Word-level Adversarial ReProgramming
Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. 2021 · 2021
Closest in time.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Closest in time.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Closest in time.
Pretrained transformers as universal computation engines
Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch. 2021 · 2021
Closest in time.
Overcoming catastrophic forgetting during domain adaptation of neural machine translation
Brian Thompson, Jeremy Gwinnup, Huda Khayrallah, Kevin Duh, and Philipp Koehn. 2019 · 2068
Closest in time.