Fetching the paper…
Reading the bibliography…
Recent research has proposed a series of specialized optimization algorithms for deep multi-task models.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Obituary: Fred jelinek
Mark Liberman · 2010
Earlier work this paper cites.
Multiple-gradient descent algorithm (mgda) for multiobjective optimization
Jean-Antoine Désidéri · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2016
Earlier work this paper cites.
Lstm: A search space odyssey
Klaus Greff, Rupesh K Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber · 2016
Earlier work this paper cites.
50 years of data science
David Donoho · 2017
Earlier work this paper cites.
Automated curriculum learning for neural networks
Alex Graves, Marc G Bellemare, Jacob Menick, Remi Munos, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich · 2018
Cited alongside, same era.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla · 2018
Cited alongside, same era.
A call for clarity in reporting bleu scores
Matt Post · 2018
Cited alongside, same era.
Multi-task learning as multi-objective optimization
Ozan Sener and Vladlen Koltun · 2018
Cited alongside, same era.
Massively multilingual neural machine translation in the wild: Findings and challenges
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Later among the works it cites.
Efficiently identifying task groupings for multi-task learning
Chris Fifty, Ehsan Amid, Zhe Zhao, Tianhe Yu, Rohan Anil, and Chelsea Finn · 2021
Later among the works it cites.
Bandits don’t follow rules: Balancing multi-facet machine translation with multi-armed bandits
Julia Kreutzer, David Vilar, and Artem Sokolov · 2021
Later among the works it cites.
Robust optimization for multilingual translation with imbalanced data
Xian Li and Hongyu Gong · 2021
Later among the works it cites.
A closer look at loss weighting in multi-task learning
Baijiong Lin, Feiyang Ye, and Yu Zhang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, et al · 2019
Cited alongside, same era.
Just pick a sign: Optimizing deep multitask models with gradient sign dropout
Zhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong, Henrik Kretzschmar, Yuning Chai, and Dragomir Anguelov · 2020
Cited alongside, same era.
In search of lost domain generalization
Ishaan Gulrajani and David Lopez-Paz · 2020
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Cited alongside, same era.
Zirui Wang, Yulia Tsvetkov, Orhan Firat, and Yuan Cao · 2020
Cited alongside, same era.
Towards impartial multi-task learning
Liyang Liu, Yi Li, Zhanghui Kuang, J Xue, Yimin Chen, Wenming Yang, Qingmin Liao, and Wayne Zhang · 2021
Later among the works it cites.
Unsupervised domain adaptation: A reality check
Kevin Musgrave, Serge Belongie, and Ser-Nam Lim · 2021
Later among the works it cites.
mslam: Massively multilingual joint pre-training for speech and text
Ankur Bapna, Colin Cherry, Yu Zhang, Ye Jia, Melvin Johnson, Yong Cheng, Simran Khanuja, Jason Riesa, and Alexis Conneau · 2022
Closest in time.
In defense of the unitary scalarization for deep multi-task learning
Vitaly Kurin, Alessandro De Palma, Ilya Kostrikov, Shimon Whiteson, and M Pawan Kumar · 2022
Closest in time.
A comparison of strategies for source-free domain adaptation
Xin Su, Yiyun Zhao, and Steven Bethard · 2022
Closest in time.