Fetching the paper…
Reading the bibliography…
Multilingual neural machine translation (MNMT) learns to translate multiple language pairs with a single model, potentially improving both the accuracy and the memory-efficiency of deployed models.
Massively multilingual neural machine translation in the wild: Findings and challenges
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, et al. 2019 · 1907
Earlier work this paper cites.
Information-type measures of difference of probability distributions and indirect observation
Imre Csiszár. 1967 · 1967
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe. 2004 · 2004
Earlier work this paper cites.
Prox-method with rate of convergence O ( 1 / t ) O(1/t) for variational inequalities with Lipschitz continuous monotone operators and smooth convex-concave saddle point problems
Arkadi Nemirovski. 2004 · 2004
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro. 2009 · 2009
Earlier work this paper cites.
Distributionally robust optimization under moment uncertainty with application to data-driven problems
Erick Delage and Yinyu Ye. 2010 · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck. 2015 · 2015
Earlier work this paper cites.
Statistics of robust optimization: A generalized empirical likelihood approach
John Duchi, Peter Glynn, and Hongseok Namkoong. 2016 · 2016
Earlier work this paper cites.
Multi-way, multilingual neural machine translation with a shared attention mechanism
Orhan Firat, Kyunghyun Cho, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Toward multilingual neural machine translation with universal encoder and decoder
Thanh-Le Ha, Jan Niehues, and Alexander Waibel. 2016 · 2016
Earlier work this paper cites.
Stochastic gradient methods for distributionally robust optimization with f f -divergences
Hongseok Namkoong and John C. Duchi. 2016 · 2016
Cited alongside, same era.
Twenty Lectures on Algorithmic Game Theory
Tim Roughgarden. 2016 · 2016
Cited alongside, same era.
Transfer learning for low-resource neural machine translation
Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight. 2016 · 2016
Cited alongside, same era.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, et al. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Data-driven robust optimization
Dimitris Bertsimas, Vishal Gupta, and Nathan Kallus. 2018 · 2018
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Variance-based regularization with convex objectives
John C. Duchi and Hongseok Namkoong. 2019 · 2019
Later among the works it cites.
Distributionally robust language modeling
Yonatan Oren, Shiori Sagawa, Tatsunori Hashimoto, and Percy Liang. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
How multilingual is multilingual bert?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Later among the works it cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fairness without demographics in repeated loss minimization
Tatsunori Hashimoto, Megha Srivastava, Hongseok Namkoong, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Rapid adaptation of neural machine translation to new languages
Graham Neubig and Junjie Hu. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
When and why are pre-trained word embeddings useful for neural machine translation?
Ye Qi, Devendra Sachan, Matthieu Felix, Sarguna Padmanabhan, and Graham Neubig. 2018 · 2018
Cited alongside, same era.
Massively multilingual neural machine translation
Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Learning models with uniform performance via distributionally robust optimization
John C. Duchi and Hongseok Namkoong. 2020 · 2020
Later among the works it cites.
Large-scale methods for distributionally robust optimization
Daniel Levy, Yair Carmon, John C. Duchi, and Aaron Sidford. 2020 · 2020
Later among the works it cites.
Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang. 2020 · 2020
Later among the works it cites.
On negative interference in multilingual language models
Zirui Wang, Zachary C Lipton, and Yulia Tsvetkov. 2020b · 2020
Later among the works it cites.
Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models
Zirui Wang, Yulia Tsvetkov, Orhan Firat, and Yuan Cao. 2021 · 2021
Closest in time.