Fetching the paper…
Reading the bibliography…
Massively multilingual models subsuming tens or even hundreds of languages pose great challenges to multi-task optimization.
Massively multilingual neural machine translation
Roee Aharoni, Melvin Johnson, and Orhan Firat · 1903
Earlier work this paper cites.
Unicoder: A universal language encoder by pre-training with multiple cross-lingual tasks
Haoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, and Ming Zhou · 1909
Earlier work this paper cites.
Balancing training for multilingual neural machine translation
Xinyi Wang, Yulia Tsvetkov, and Graham Neubig · 2004
Earlier work this paper cites.
Tags for identifying languages
Addison Phillips and Mark Davis · 2006
Earlier work this paper cites.
Large scale parallel document mining for machine translation
Jakob Uszkoreit, Jay Ponte, Ashok Popat, and Moshe Dubiner · 2010
Earlier work this paper cites.
A convex formulation for learning task relationships in multi-task learning
Yu Zhang and Dit-Yan Yeung · 2012
Earlier work this paper cites.
On handling negative transfer and imbalanced distributions in multiple source transfer learning
Liang Ge, Jing Gao, Hung Ngo, Kang Li, and Aidong Zhang · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Multi-way, multilingual neural machine translation with a shared attention mechanism
Orhan Firat, Kyunghyun Cho, and Yoshua Bengio · 2016
Earlier work this paper cites.
Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell · 2016
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc V Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, et al · 2017
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji · 2017
Earlier work this paper cites.
An overview of multi-task learning in deep neural networks
Sebastian Ruder · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Multilingual neural machine translation with task-specific attention
Graeme W. Blackwood, Miguel Ballesteros, and Todd Ward · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Adapting auxiliary losses using gradient similarity
Yunshu Du, Wojciech M Czarnecki, Siddhant M Jayakumar, Razvan Pascanu, and Balaji Lakshminarayanan · 2018
Cited alongside, same era.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla · 2018
Cited alongside, same era.
Taku Kudo and John Richardson · 2018
Cited alongside, same era.
Universal dependencies 2.2
Joakim Nivre, Mitchell Abrams, Željko Agić, Lars Ahrenberg, Lene Antonsen, Maria Jesus Aranzabe, Gashaw Arutie, Masayuki Asahara, Luma Ateyah, Mohammed Attia, et al · 2018
Pareto multi-task learning
Xi Lin, Hui-Ling Zhen, Zhenhua Li, Qing-Fu Zhang, and Sam Kwong · 2019
Later among the works it cites.
Polyglot contextual representations improve crosslingual transfer
Phoebe Mulcaire, Jungo Kasai, and Noah A Smith · 2019
Later among the works it cites.
How multilingual is multilingual bert?
Telmo Pires, Eva Schlinger, and Dan Garrette · 2019
Later among the works it cites.
Routing networks and the challenges of modular and compositional computation
Clemens Rosenbaum, Ignacio Cases, Matthew Riemer, and Tim Klinger · 2019
Later among the works it cites.
Ray interference: a source of plateaus in deep reinforcement learning
Tom Schaul, Diana Borsa, Joseph Modayil, and Razvan Pascanu · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Parameter sharing methods for multilingual self-attentional translation models
Devendra Sachan and Graham Neubig · 2018
Cited alongside, same era.
Multi-task learning as multi-objective optimization
Ozan Sener and Vladlen Koltun · 2018
Cited alongside, same era.
Towards more reliable transfer learning
Zirui Wang and Jaime Carbonell · 2018
Cited alongside, same era.
Taskonomy: Disentangling task transfer learning
Amir R Zamir, Alexander Sax, William Shen, Leonidas J Guibas, Jitendra Malik, and Silvio Savarese · 2018
Cited alongside, same era.
Massively multilingual neural machine translation in the wild: Findings and challenges
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, et al · 2019
Cited alongside, same era.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Multilingual NMT with a language-independent attention bridge
Raúl Vázquez, Alessandro Raganato, Jörg Tiedemann, and Mathias Creutz · 2019
Later among the works it cites.
Characterizing and avoiding negative transfer
Zirui Wang, Zihang Dai, Barnabás Póczos, and Jaime Carbonell · 2019
Later among the works it cites.
Beto, bentz, becas: The surprising cross-lingual effectiveness of bert
Shijie Wu and Mark Dredze · 2019
Later among the works it cites.
Emerging cross-lingual structure in pretrained language models
Shijie Wu, Alexis Conneau, Haoran Li, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Later among the works it cites.
Multilingual alignment of contextual word representations
Steven Cao, Nikita Kitaev, and Dan Klein · 2020
Closest in time.
Carlos Escolano, Marta R Costa-jussà, José AR Fonollosa, and Mikel Artetxe · 2020
Closest in time.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson · 2020
Closest in time.
Cross-lingual ability of multilingual bert: An empirical study
K Karthikeyan, Zihan Wang, Stephen Mayhew, and Dan Roth · 2020
Closest in time.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Closest in time.
Evaluating the cross-lingual effectiveness of massively multilingual neural machine translation
Aditya Siddhant, Melvin Johnson, Henry Tsai, Naveen Ari, Jason Riesa, Ankur Bapna, Orhan Firat, and Karthik Raman · 2020
Closest in time.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Closest in time.