Fetching the paper…
Reading the bibliography…
Multilingual models jointly pretrained on multiple languages have achieved remarkable performance on various multilingual downstream tasks.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky et al · 2009
Earlier work this paper cites.
A convex formulation for learning task relationships in multi-task learning
Yu Zhang and Dit-Yan Yeung · 2010
Earlier work this paper cites.
Learning with whom to share in multi-task feature learning
Zhuoliang Kang, Kristen Grauman, and Fei Sha · 2011
Earlier work this paper cites.
Novel dataset for fine-grained image categorization: Stanford dogs
Aditya Khosla, Nityananda Jayadevaprakash, Bangpeng Yao, and Fei-Fei Li · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie · 2011
Earlier work this paper cites.
Learning task grouping and overlap in multi-task learning
Abhishek Kumar and Hal Daumé III · 2012
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi · 2013
Earlier work this paper cites.
Tiny imagenet visual recognition challenge
Ya Le and Xuan Yang · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Asymmetric multi-task learning based on task relatedness and loss
Giwoong Lee, Eunho Yang, and Sung Hwang · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich · 2018
Earlier work this paper cites.
Xnli: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov · 2018
Cited alongside, same era.
On first-order meta-learning algorithms
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Cited alongside, same era.
Multi-task learning as multi-objective optimization
Ozan Sener and Vladlen Koltun · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Cited alongside, same era.
Massively multilingual neural machine translation in the wild: Findings and challenges
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, et al · 2019
Cited alongside, same era.
Look-ahead meta learning for continual learning
Gunshi Gupta, Karmesh Yadav, and Liam Paull · 2020
Later among the works it cites.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer · 2020
Later among the works it cites.
Zero-shot cross-lingual transfer with meta learning
Farhad Nooralahzadeh, Giannis Bekoulis, Johannes Bjerva, and Isabelle Augenstein · 2020
Later among the works it cites.
On negative interference in multilingual language models
Zirui Wang, Zachary C Lipton, and Yulia Tsvetkov · 2020
Later among the works it cites.
Ccnet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Édouard Grave · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multilingual alignment of contextual word representations
Steven Cao, Nikita Kitaev, and Dan Klein · 2019
Cited alongside, same era.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Pareto multi-task learning
Xi Lin, Hui-Ling Zhen, Zhenhua Li, Qing-Fu Zhang, and Sam Kwong · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Learning to learn without forgetting by maximizing transfer and minimizing interference
Matthew Riemer, Ignacio Cases, Robert Ajemian, Miao Liu, Irina Rish, Yuhai Tu, , and Gerald Tesauro · 2019
Cited alongside, same era.
Characterizing and avoiding negative transfer
Zirui Wang, Zihang Dai, Barnabás Póczos, and Jaime G. Carbonell · 2019
Cited alongside, same era.
Thomas Wolf, Julien Chaumond, Lysandre Debut, Victor Sanh, Clement Delangue, Anthony Moi, Pierric Cistac, Morgan Funtowicz, Joe Davison, Sam Shleifer, et al · 2020
Later among the works it cites.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Later among the works it cites.
Explicit alignment objectives for multilingual bidirectional encoders
Junjie Hu, Melvin Johnson, Orhan Firat, Aditya Siddhant, and Graham Neubig · 2021
Closest in time.
Learning to perturb word embeddings for out-of-distribution QA
Seanie Lee, Minki Kang, Juho Lee, and Sung Ju Hwang · 2021
Closest in time.
Robust optimization for multilingual translation with imbalanced data
Xian Li and Hongyu Gong · 2021
Closest in time.
Multilingual bert post-pretraining alignment
Lin Pan, Chung-Wei Hang, Haode Qi, Abhishek Shah, Saloni Potdar, and Mo Yu · 2021
Closest in time.
Conditionally adaptive multi-task learning: Improving transfer learning in NLP using fewer parameters & less data
Jonathan Pilault, Amine El hattami, and Christopher Pal · 2021
Closest in time.
Large-scale meta-learning with continual trajectory shifting
Jaewoong Shin, Haebeom Lee, Boqing Gong, and Sung Ju Hwang · 2021
Closest in time.
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David Barrett, and Soham De · 2021
Closest in time.
Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models
Zirui Wang, Yulia Tsvetkov, Orhan Firat, and Yuan Cao · 2021
Closest in time.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel · 2021
Closest in time.