Fetching the paper…
Reading the bibliography…
When training multilingual machine translation (MT) models that can translate to/from multiple languages, we are faced with imbalanced training sets: some languages have much more training data than others.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams. 1992 · 1992
Earlier work this paper cites.
Multilingual and crosslingual speech recognition
Tanja Schultz and Alex Waibel. 1998 · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
An overview of bilevel optimization
Benoît Colson, Patrice Marcotte, and Gilles Savard. 2007 · 2007
Earlier work this paper cites.
Polylingual topic models
David M. Mimno, Hanna M. Wallach, Jason Naradowsky, David A. Smith, and Andrew McCallum. 2009 · 2009
Earlier work this paper cites.
Cross language text classification by model translation and semi-supervised learning
Lei Shi, Rada Mihalcea, and Mingjun Tian. 2010 · 2010
Earlier work this paper cites.
Inducing crosslingual distributed representations of words
Alexandre Klementiev, Ivan Titov, and Binod Bhattarai. 2012 · 2012
Earlier work this paper cites.
Token and type constraints for cross-lingual part-of-speech tagging
Oscar Täckström, Dipanjan Das, Slav Petrov, Ryan McDonald, and Joakim Nivre. 2013 · 2013
Earlier work this paper cites.
Multi-task learning for multiple language translation
Daxiang Dong, Hua Wu, Wei He, Dianhai Yu, and Haifeng Wang. 2015 · 2015
Earlier work this paper cites.
Many languages, one parser
Waleed Ammar, George Mulcaire, Miguel Ballesteros, Chris Dyer, and Noah A Smith. 2016 · 2016
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson et al. 2016 · 2016
Earlier work this paper cites.
Multilingual part-of-speech tagging with bidirectional long short-term memory models and auxiliary loss
Barbara Plank, Anders Søgaard, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Transfer learning for low resource neural machine translation
Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Cited alongside, same era.
Learning language representations for typology prediction
Chaitanya Malaviya, Graham Neubig, and Patrick Littell. 2017 · 2017
Cited alongside, same era.
Continuous multilinguality with language vectors
Robert Östling and Jörg Tiedemann. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martínez-Rubio, Mark Schmidt, and Frank Wood. 2018 · 2018
Cited alongside, same era.
Meta-learning for low-resource neural machine translation
Jiatao Gu, Yong Wang, Yun Chen, Victor O. K. Li, and Kyunghyun Cho. 2018 · 2018
Learning to reweight examples for robust deep learning
Mengye Ren, Wenyuan Zeng, Bin Yang, and Raquel Urtasun. 2018 · 2018
Later among the works it cites.
Three strategies to improve one-to-many multilingual translation
Yining Wang, Jiajun Zhang, Feifei Zhai, Jingfang Xu, and Chengqing Zong. 2018 · 2018
Later among the works it cites.
Neural cross-lingual named entity recognition with minimal resources
Jiateng Xie, Zhilin Yang, Graham Neubig, Noah A. Smith, and Jaime Carbonell. 2018 · 2018
Later among the works it cites.
Adaptive knowledge sharing in multi-task learning: Improving low-resource neural machine translation
Poorya Zaremoodi, Wray L. Buntine, and Gholamreza Haffari. 2018 · 2018
Later among the works it cites.
Massively multilingual neural machine translation
Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019 · 2019
Later among the works it cites.
Massively multilingual neural machine translation in the wild: Findings and challenges
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla. 2018 · 2018
Cited alongside, same era.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Rapid adaptation of neural machine translation to new languages
Graham Neubig and Junjie Hu. 2018 · 2018
Cited alongside, same era.
Transfer learning across low-resource, related languages for neural machine translation
Toan Q. Nguyen and David Chiang. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
When and why are pre-trained word embeddings useful for neural machine translation?
Ye Qi, Devendra Singh Sachan, Matthieu Felix, Sarguna Padmanabhan, and Graham Neubig. 2018 · 2018
Cited alongside, same era.
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, Wolfgang Macherey, Zhifeng Chen, and Yonghui Wu. 2019 · 2019
Later among the works it cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Later among the works it cites.
Choosing transfer languages for cross-lingual learning
Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, and Graham Neubig. 2019 · 2019
Later among the works it cites.
Transformers without tears: Improving the normalization of self-attention
Toan Q. Nguyen and Julián Salazar. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Target conditioned sampling: Optimizing data selection for multilingual neural machine translation
Xinyi Wang and Graham Neubig. 2019 · 2019
Later among the works it cites.
Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT
Shijie Wu and Mark Dredze. 2019 · 2019
Later among the works it cites.