Multitask learning
R. Caruana · 1997
Earlier work this paper cites.
Balancing training for multilingual neural machine translation
Original
X. Wang, Y. Tsvetkov, and G. Neubig · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
On negative interference in multilingual models: Findings and a meta-learning treatment
Original
Z. Wang, Z. C. Lipton, and Y. Tsvetkov · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. Berg · 2011
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Original
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Original
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Real-time analysis and visualization of the yfcc100m dataset
S. Kalkowski, C. Schulze, A. Dengel, and D. Borth · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Earlier work this paper cites.
Layer normalization
Original
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
A joint many-task model: Growing a neural network for multiple nlp tasks
Original
K. Hashimoto, C. Xiong, Y. Tsuruoka, and R. Socher · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Original
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Earlier work this paper cites.
Adversarial multi-task learning for text classification
Original
P. Liu, X. Qiu, and X. Huang · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Original
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Z. Chen, V. Badrinarayanan, C.-Y. Lee, and A. Rabinovich · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Dynamic task prioritization for multitask learning
M. Guo, A. Haque, D.-A. Huang, S. Yeung, and L. Fei-Fei · 2018
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
A. Kendall, Y. Gal, and R. Cipolla · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
N. Shazeer, Y. Cheng, N. Parmar, D. Tran, A. Vaswani, P. Koanantakool, P. Hawkins, H. Lee, M. Hong, C. Young, et al · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Original
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2018
Earlier work this paper cites.
Massively multilingual neural machine translation in the wild: Findings and challenges
Original
N. Arivazhagan, A. Bapna, O. Firat, D. Lepikhin, M. Johnson, M. Krikun, M. X. Chen, Y. Cao, G. Foster, C. Cherry, et al · 2019
Earlier work this paper cites.