Fetching the paper…
Reading the bibliography…
Linear scalarization, i.e., combining all loss functions by a weighted sum, has been the default choice in the literature of multi-task learning (MTL) since its inception.
The approximation of one matrix by another of lower rank
C. Eckart and G. Young · 1936
Earlier work this paper cites.
Optimality and non-scalar-valued performance criteria
L. Zadeh · 1963
Earlier work this paper cites.
Directional convexity and the maximum principle for discrete systems
J. M. Holtzman and H. Halkin · 1966
Earlier work this paper cites.
Three methods for determining pareto-optimal solutions of multiple-objective problems
J. G. Lin · 1976
Earlier work this paper cites.
Fundamentals of statistical signal processing: estimation theory
S. M. Kay · 1993
Earlier work this paper cites.
Multitask learning
R. Caruana · 1997
Earlier work this paper cites.
Convex analysis , volume 11
R. T. Rockafellar · 1997
Earlier work this paper cites.
Matrix iterative analysis , volume 27
R. S. Varga · 1999
Earlier work this paper cites.
Steepest descent methods for multicriteria optimization
J. Fliege and B. F. Svaiter · 2000
Earlier work this paper cites.
Locally weighted projection regression: An o (n) algorithm for incremental real time learning in high dimensional space
S. Vijayakumar and S. Schaal · 2000
Earlier work this paper cites.
Convex optimization
S. Boyd, S. P. Boyd, and L. Vandenberghe · 2004
Earlier work this paper cites.
Efficient algorithms for online decision problems
A. Kalai and S. Vempala · 2005
Earlier work this paper cites.
Bounds for linear multi-task learning
A. Maurer · 2006
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
R. Collobert and J. Weston · 2008
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction , volume 2
T. Hastie, R. Tibshirani, J. H. Friedman, and J. H. Friedman · 2009
Earlier work this paper cites.
Multiple-gradient descent algorithm (mgda) for multiobjective optimization
J.-A. Désidéri · 2012
Earlier work this paper cites.
Matrix analysis , volume 169
R. Bhatia · 2013
Earlier work this paper cites.
Linear and nonlinear functional analysis with applications , volume 130
P. G. Ciarlet · 2013
Cited alongside, same era.
The benefit of multitask representation learning
A. Maurer, M. Pontil, and B. Romera-Paredes · 2016
Cited alongside, same era.
Cross-stitch networks for multi-task learning
I. Misra, A. Shrivastava, A. Gupta, and M. Hebert · 2016
Cited alongside, same era.
An overview of multi-task learning in deep neural networks
S. Ruder · 2017
Cited alongside, same era.
Multi-task learning as multi-objective optimization
O. Sener and V. Koltun · 2018
Cited alongside, same era.
An overview of multi-task learning
Y. Zhang and Q. Yang · 2018
Cited alongside, same era.
Provable meta-learning of linear representations
N. Tripuraneni, C. Jin, and M. Jordan · 2021
Later among the works it cites.
Multi-task learning for dense prediction tasks: A survey
S. Vandenhende, S. Georgoulis, W. Van Gansbeke, M. Proesmans, D. Dai, and L. Van Gool · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2021
Later among the works it cites.
A survey on multi-task learning
Y. Zhang and Q. Yang · 2021
Later among the works it cites.
Mitigating gradient bias in multi-objective learning: A provably convergent approach
H. D. Fernando, H. Shen, M. Liu, S. Chaudhury, K. Murugesan, and T. Chen · 2022
Later among the works it cites.
In defense of the unitary scalarization for deep multi-task learning
V. Kurin, A. De Palma, I. Kostrikov, S. Whiteson, and P. K. Mudigonda · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pareto multi-task learning
X. Lin, H.-L. Zhen, Z. Li, Q.-F. Zhang, and S. Kwong · 2019
Cited alongside, same era.
Just pick a sign: Optimizing deep multitask models with gradient sign dropout
Z. Chen, J. Ngiam, Y. Huang, T. Luong, H. Kretzschmar, Y. Chai, and D. Anguelov · 2020
Cited alongside, same era.
Multi-task learning with deep neural networks: A survey
M. Crawshaw · 2020
Cited alongside, same era.
Controllable pareto multi-task learning
X. Lin, Z. Yang, Q. Zhang, and S. Kwong · 2020
Cited alongside, same era.
Efficient continuous pareto exploration in multi-task learning
P. Ma, T. Du, and W. Matusik · 2020
Cited alongside, same era.
Multi-task learning with user preferences: Gradient descent with controlled ascent in pareto optimization
D. Mahapatra and V. Rajan · 2020
Cited alongside, same era.
Later among the works it cites.
Reasonable effectiveness of random weighting: A litmus test for multi-task learning
B. Lin, Y. Feiyang, Y. Zhang, and I. Tsang · 2022
Later among the works it cites.
A multi-objective/multi-task learning framework induced by pareto stationarity
M. Momma, C. Dong, and J. Liu · 2022
Later among the works it cites.
Multi-task learning as a bargaining game
A. Navon, A. Shamsian, I. Achituve, H. Maron, K. Kawaguchi, G. Chechik, and E. Fetaya · 2022
Later among the works it cites.
Can small heads help? understanding and improving multi-task generalization
Y. Wang, Z. Zhao, B. Dai, C. Fifty, D. Lin, L. Hong, L. Wei, and E. H. Chi · 2022
Later among the works it cites.
Do current multi-task optimization methods in deep learning even help?
D. Xin, B. Ghorbani, J. Gilmer, A. Garg, and O. Firat · 2022
Later among the works it cites.
Pareto navigation gradient descent: a first-order algorithm for optimization in pareto set
M. Ye and Q. Liu · 2022
Later among the works it cites.
Inherent tradeoffs in learning fair representations
H. Zhao and G. J. Gordon · 2022
Later among the works it cites.
On the convergence of stochastic multi-objective gradient manipulation and beyond
S. Zhou, W. Zhang, J. Jiang, W. Zhong, J. Gu, and W. Zhu · 2022
Later among the works it cites.
Three-way trade-off in multi-objective learning: Optimization, generalization and conflict-avoidance
L. Chen, H. Fernando, Y. Ying, and T. Chen · 2023
Closest in time.
Fair and optimal classification via post-processing
R. Xian, L. Yin, and H. Zhao · 2023
Closest in time.