Fetching the paper…
Reading the bibliography…
Merging multiple expert models offers a promising approach for performing multi-task learning without accessing their original data.
The MNIST database of handwritten digits
LeCun, Y · 1998
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., Ng, A. Y., et al · 2011
Earlier work this paper cites.
The german traffic sign recognition benchmark: a multi-class classification competition
Stallkamp, J., Schlipsing, M., Salmen, J., and Igel, C · 2011
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A · 2014
Earlier work this paper cites.
Sun database: Exploring a large collection of scene categories
Xiao, J., Ehinger, K. A., Hays, J., Torralba, A., and Oliva, A · 2016
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
Cheng, G., Han, J., and Lu, X · 2017
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science
Vershynin, R · 2018
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Helber, P., Bischke, B., Dengel, A., and Borth, D · 2019
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Earlier work this paper cites.
Versatile black-box optimization
Liu, J., Moreau, A., Preuss, M., Rapin, J., Roziere, B., Teytaud, F., and Teytaud, O · 2020
Earlier work this paper cites.
Adashare: Learning what to share for efficient deep multi-task learning
Sun, X., Panda, R., Feris, R., and Saenko, K · 2020
Earlier work this paper cites.
Gradient surgery for multi-task learning
Yu, T., Kumar, S., Gupta, A., Levine, S., Hausman, K., and Finn, C · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
Gradient projection memory for continual learning
Saha, G., Garg, I., and Roy, K · 2021
Earlier work this paper cites.
Metagpt: Merging large language models using model exclusive task arithmetic
Zhou, Y., Song, L., Wang, B., and Chen, W · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al · 2022
Cited alongside, same era.
Merging models with fisher-weighted averaging
Matena, M. S. and Raffel, C. A · 2022
Cited alongside, same era.
Recon: Reducing conflicting gradients from the root for multi-task learning
Shi, G., Li, Q., Zhang, W., Chen, J., and Wu, X.-M · 2022
Cited alongside, same era.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., and Schmidt, L · 2022
Cited alongside, same era.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2023
Cited alongside, same era.
Forkmerge: Mitigating negative transfer in auxiliary-task learning
Task singular vectors: Reducing task interference in model merging
Gargiulo, A. A., Crisostomi, D., Bucarelli, M. S., Scardapane, S., Silvestri, F., and Rodolà, E · 2024
Later among the works it cites.
Harmony in diversity: Merging neural networks with canonical correlation analysis
Horoi, S., Camacho, A. M. O., Belilovsky, E., and Wolf, G · 2024
Later among the works it cites.
EMR-merging: Tuning-free high-performance model merging
Huang, C., Ye, P., Chen, T., He, T., Yue, X., and Ouyang, W · 2024
Later among the works it cites.
Twin-merging: Dynamic integration of modular expertise in model merging
Lu, Z., Fan, C., Wei, W., Qu, X., Chen, D., and Cheng, Y · 2024
Later among the works it cites.
How to weight multitask finetuning? fast previews via bayesian model-merging
Maldonado, H. M., Möllenhoff, T., Daheim, N., Gurevych, I., and Khan, M. E · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiang, J., Chen, B., Pan, J., Wang, X., Liu, D., Long, M., et al · 2023
Cited alongside, same era.
Li, W., Peng, Y., Zhang, M., Ding, L., Hu, H., and Shen, L · 2023
Cited alongside, same era.
Task arithmetic in the tangent space: Improved editing of pre-trained models
Ortiz-Jimenez, G., Favero, A., and Frossard, P · 2023
Cited alongside, same era.
Concrete subspace learning based interference elimination for multi-task model fusion
Tang, A., Shen, L., Luo, Y., Ding, L., Hu, H., Du, B., and Tao, D · 2023
Cited alongside, same era.
Ties-merging: Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C. A., and Bansal, M · 2023
Cited alongside, same era.
Local superior soups: A catalyst for model merging in cross-silo federated learning
Chen, M., Jiang, M., Zhang, X., Dou, Q., Wang, Z., and Li, X · 2024
Cited alongside, same era.
Revisiting weight averaging for model merging
Choi, J., Kim, D., Lee, C., and Hong, S · 2024
Cited alongside, same era.
Learning to route among specialized experts for zero-shot generalization
Muqeeth, M., Liu, H., Liu, Y., and Raffel, C · 2024
Later among the works it cites.
Efficient and effective weight-ensembling mixture of experts for multi-task model merging
Shen, L., Tang, A., Yang, E., Guo, G., Luo, Y., Zhang, L., Cao, X., Du, B., and Tao, D · 2024
Later among the works it cites.
Zipit! merging models from different tasks without training
Stoica, G., Bolya, D., Bjorner, J., Ramesh, P., Hearn, T., and Hoffman, J · 2024
Later among the works it cites.
Task groupings regularization: Data-free meta-learning with heterogeneous pre-trained models
Wei, Y., Hu, Z., Shen, L., Wang, Z., Li, Y., Yuan, C., and Tao, D · 2024
Later among the works it cites.
Multi-task model merging via adaptive weight disentanglement
Xiong, F., Cheng, R., Chen, W., Zhang, Z., Guo, Y., Yuan, C., and Xu, R · 2024
Later among the works it cites.
Language models are super mario: Absorbing abilities from homologous models as a free lunch
Yu, L., Yu, B., Yu, H., Huang, F., and Li, Y · 2024
Later among the works it cites.
Model merging with svd to tie the knots
Stoica, G., Ramesh, P., Ecsedi, B., Choshen, L., and Hoffman, J · 2025
Closest in time.
Sun, W., Li, Q., Wang, W., Geng, Y.-a., and Li, B · 2025
Closest in time.
Open-vocabulary customization from CLIP via data-free knowledge distillation
Wei, Y., Hu, Z., Shen, L., Wang, Z., Yuan, C., and Tao, D · 2025
Closest in time.
Zhang, Q., Qi, Y., Tang, X., Yuan, R., Lin, X., Zhang, K., and Yuan, C · 2025
Closest in time.