Fetching the paper…
Reading the bibliography…
Model merging-based multitask learning (MTL) offers a promising approach for performing MTL by merging multiple expert models without requiring access to raw training data.
R. A. Fisher, “On the mathematical foundations of theoretical statistics,”
1922
Earlier work this paper cites.
J. Ma, Z. Zhao, X. Yi, J. Chen, L. Hong, and E. H. Chi, “Modeling task relationships in multi-task learning with multi-gate mixture-of-experts,” in
1939
Earlier work this paper cites.
R. Caruana, “Multitask learning,”
1997
Earlier work this paper cites.
Y. LeCun, “The mnist database of handwritten digits,”
1998
Earlier work this paper cites.
R. Collobert and J. Weston, “A unified architecture for natural language processing: Deep neural networks with multitask learning,” in
2008
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.”
2008
Earlier work this paper cites.
S. J. Pan and Q. Yang, “A survey on transfer learning,”
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
2009
Earlier work this paper cites.
N. Yuval, “Reading digits in natural images with unsupervised feature learning,” in
2011
Earlier work this paper cites.
J. Stallkamp, M. Schlipsing, J. Salmen, and C. Igel, “The german traffic sign recognition benchmark: a multi-class classification competition,” in
2011
Earlier work this paper cites.
J. Krause, M. Stark, J. Deng, and L. Fei-Fei, “3d object representations for fine-grained categorization,” in
2013
Earlier work this paper cites.
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi, “Describing textures in the wild,” in
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
D. Dong, H. Wu, W. He, D. Yu, and H. Wang, “Multi-task learning for multiple language translation,” in
2015
Earlier work this paper cites.
X. Zhang, J. Zhao, and Y. LeCun, “Character-level convolutional networks for text classification,”
2015
Earlier work this paper cites.
J. Xiao, K. A. Ehinger, J. Hays, A. Torralba, and A. Oliva, “Sun database: Exploring a large collection of scene categories,”
2016
Earlier work this paper cites.
I. Misra, A. Shrivastava, A. Gupta, and M. Hebert, “Cross-stitch networks for multi-task learning,” in
2016
Earlier work this paper cites.
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in
2017
Earlier work this paper cites.
G. Cheng, J. Han, and X. Lu, “Remote sensing image scene classification: Benchmark and state of the art,”
2017
Earlier work this paper cites.
G. Hadash, O. S. Shalom, and R. Osadchy, “Rank and rate: multi-task learning for recommender systems,” in
2018
Earlier work this paper cites.
Z. Chen, V. Badrinarayanan, C.-Y. Lee, and A. Rabinovich, “Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks,” in
2018
Earlier work this paper cites.
O. Sener and V. Koltun, “Multi-task learning as multi-objective optimization,” in
2018
Earlier work this paper cites.
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson, “Averaging weights leads to wider optima and better generalization,”
2018
Earlier work this paper cites.
A. Jacot, F. Gabriel, and C. Hongler, “Neural tangent kernel: Convergence and generalization in neural networks,”
2018
Earlier work this paper cites.
A. Kendall, Y. Gal, and R. Cipolla, “Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,” in
2018
Earlier work this paper cites.
S. Liu, E. Johns, and A. J. Davison, “End-to-end multi-task learning with attention,” in
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” in
2019
Earlier work this paper cites.
P. Helber, B. Bischke, A. Dengel, and D. Borth, “Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,”
2019
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in
2019
Earlier work this paper cites.
Z. Chen, J. Ngiam, Y. Huang, T. Luong, H. Kretzschmar, Y. Chai, and D. Anguelov, “Just pick a sign: Optimizing deep multitask models with gradient sign dropout,” in
2020
Earlier work this paper cites.
H. Tang, J. Liu, M. Zhao, and X. Gong, “Progressive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations,” in
2020
Cited alongside, same era.
M. Ravanelli, J. Zhong, S. Pascual, P. Swietojanski, J. Monteiro, J. Trmal, and Y. Bengio, “Multi-task self-supervised learning for robust speech recognition,” in
2020
Cited alongside, same era.
S. P. Singh and M. Jaggi, “Model fusion via optimal transport,”
2020
Cited alongside, same era.
X. Sun, R. Panda, R. Feris, and K. Saenko, “Adashare: Learning what to share for efficient deep multi-task learning,”
2020
Cited alongside, same era.
Y. Gao, H. Bai, Z. Jie, J. Ma, K. Jia, and W. Liu, “Mtl-nas: Task-agnostic neural architecture search towards general-purpose multi-task learning,” in
2020
Cited alongside, same era.
C. Wu, T. Wang, Y. Ge, Z. Lu, R. Zhou, Y. Shan, and P. Luo, “pi-tuning: Transferring multimodal foundation models with optimal multi-task interpolation,” in
2023
Later among the works it cites.
A. Panigrahi, N. Saunshi, H. Zhao, and S. Arora, “Task-specific skill localization in fine-tuned language models,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Niu, J. Wu, Y. Zhang, Z. Wen, Y. Chen, P. Zhao, and M. Tan, “Towards stable test-time adaptation in dynamic wild world,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Yu, S. Kumar, A. Gupta, S. Levine, K. Hausman, and C. Finn, “Gradient surgery for multi-task learning,”
2020
Cited alongside, same era.
W.-H. Li and H. Bilen, “Knowledge distillation for multi-task learning,” in
2020
Cited alongside, same era.
S. Vandenhende, S. Georgoulis, W. Van Gansbeke, M. Proesmans, D. Dai, and L. Van Gool, “Multi-task learning for dense prediction tasks: A survey,”
2021
Cited alongside, same era.
S. Chen, Y. Zhang, and Q. Yang, “Multi-task learning in natural language processing: An overview,”
2021
Cited alongside, same era.
B. Liu, X. Liu, X. Jin, P. Stone, and Q. Liu, “Conflict-averse gradient descent for multi-task learning,”
2021
Cited alongside, same era.
X. Cai, J. Yuan, R. Zheng, L. Huang, and K. Church, “Speech emotion recognition with multi-task learning.” in
2021
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,”
2021
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Ainsworth, J. Hayase, and S. Srinivasa, “Git re-basin: Merging models modulo permutation symmetries,” in
2023
Later among the works it cites.
K. Jordan, H. Sedghi, O. Saukh, R. Entezari, and B. Neyshabur, “REPAIR: renormalizing permuted activations for interpolation repair,” in
2023
Later among the works it cites.
Y. Hu, R. Xian, Q. Wu, Q. Fan, L. Yin, and H. Zhao, “Revisiting scalarization in multi-task learning: A theoretical perspective,” in
2023
Later among the works it cites.
E. Yang, L. Shen, Z. Wang, G. Guo, X. Chen, X. Wang, and D. Tao, “Representation surgery for multi-task model merging,”
2024
Closest in time.
G. Stoica, D. Bolya, J. Bjorner, T. Hearn, and J. Hoffman, “Zipit! merging models from different tasks without training,”
2024
Closest in time.
E. Yang, Z. Wang, L. Shen, S. Liu, G. Guo, X. Wang, and D. Tao, “Adamerging: Adaptive model merging for multi-task learning,”
2024
Closest in time.
Z. Xu, K. Yuan, H. Wang, Y. Wang, M. Song, and J. Song, “Training-free pretrained model merging,” in
2024
Closest in time.
A. Tang, L. Shen, Y. Luo, Y. Zhan, H. Hu, B. Du, Y. Chen, and D. Tao, “Parameter efficient multi-task model fusion with partial linearization,”
2024
Closest in time.
L. Yu, B. Yu, H. Yu, F. Huang, and Y. Li, “Language models are super mario: Absorbing abilities from homologous models as a free lunch,”
2024
Closest in time.
K. Wang, N. Dimitriadis, G. Ortiz-Jimenez, F. Fleuret, and P. Frossard, “Localizing task information for improved model merging and compression,”
2024
Closest in time.
C. Huang, P. Ye, T. Chen, T. He, X. Yue, and W. Ouyang, “Emr-merging: Tuning-free high-performance model merging,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
N. Daheim, T. Möllenhoff, E. M. Ponti, I. Gurevych, and M. E. Khan, “Model merging by uncertainty-based gradient matching,” 2024
2024
Closest in time.
N. Daheim, T. Möllenhoff, E. M. Ponti, I. Gurevych, and M. E. Khan, “Model merging by uncertainty-based gradient matching,”
2024
Closest in time.
2024
Closest in time.
G. Du, J. Lee, J. Li, R. Jiang, Y. Guo, S. Yu, H. Liu, S. K. Goh, H.-K. Tang, D. He
2024
Closest in time.
M. Zimmer, C. Spiegel, and S. Pokutta, “Sparse model soups: A recipe for improved pruning via model averaging,”
2024
Closest in time.
2024
Closest in time.
A. Tang, L. Shen, Y. Luo, N. Yin, L. Zhang, and D. Tao, “Merging multi-task models via weight-ensembling mixture of experts,”
2024
Closest in time.
M. Muqeeth, H. Liu, Y. Liu, and C. Raffel, “Learning to route among specialized experts for zero-shot generalization,”
2024
Closest in time.
P. Li, Z. Zhang, P. Yadav, Y.-L. Sung, Y. Cheng, M. Bansal, and T. Chen, “Merge, then compress: Demystify efficient smoe with hints from its routing policy,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Y. Liu, A. Golatkar, and S. Soatto, “Tangent transformers for composition, privacy and removal,”
2024
Closest in time.