Fetching the paper…
Reading the bibliography…
Multi-task model merging offers an efficient solution for integrating knowledge from multiple fine-tuned models, mitigating the significant computational and storage demands associated with multi-task training.
Theory of statistical estimation
Ronald Aylmer Fisher · 1925
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Baolin Wu, Andrew Y Ng, et al · 2011
Earlier work this paper cites.
The german traffic sign recognition benchmark: A multi-class classification competition
Johannes Stallkamp, Marc Schlipsing, Jan Salmen, and Christian Igel · 2011
Earlier work this paper cites.
3d object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
All of statistics: A concise course in statistical inference, 2013
L Wasserman · 2013
Earlier work this paper cites.
Describing textures in the wild
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi · 2014
Earlier work this paper cites.
Sun database: Exploring a large collection of scene categories
Jianxiong Xiao, Krista A Ehinger, James Hays, Antonio Torralba, and Aude Oliva · 2016
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
Gong Cheng, Junwei Han, and Xiaoqiang Lu · 2017
Earlier work this paper cites.
Adversarial multi-task learning for text classification
Pengfei Liu, Xipeng Qiu, and Xuanjing Huang · 2017
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2017
Earlier work this paper cites.
GradNorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi · 2018
Earlier work this paper cites.
Multi-task learning as multi-objective optimization
Ozan Sener and Vladlen Koltun · 2018
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth · 2019
Earlier work this paper cites.
End-to-end multi-task learning with attention
Shikun Liu, Edward Johns, and Andrew J. Davison · 2019
Cited alongside, same era.
Just pick a sign: Optimizing deep multitask models with gradient sign dropout
Zhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong, Henrik Kretzschmar, Yuning Chai, and Dragomir Anguelov · 2020
Cited alongside, same era.
Analyzing redundancy in pretrained transformer models
Fahim Dalvi, Hassan Sajjad, Nadir Durrani, and Yonatan Belinkov · 2020
Cited alongside, same era.
Mtl-nas: Task-agnostic neural architecture search towards general-purpose multi-task learning
Yuan Gao, Haoping Bai, Zequn Jie, Jiayi Ma, Kui Jia, and Wei Liu · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Rotograd: Gradient homogenization in multitask learning
Adrián Javaloy and Isabel Valera · 2022
Later among the works it cites.
Auto-lambda: Disentangling dynamic task relationships
Shikun Liu, Stephen James, Andrew Davison, and Edward Johns · 2022
Later among the works it cites.
Merging models with fisher-weighted averaging
Michael S Matena and Colin A Raffel · 2022
Later among the works it cites.
Multi-task learning as a bargaining game
Aviv Navon, Aviv Shamsian, Idan Achituve, Haggai Maron, Kenji Kawaguchi, Gal Chechik, and Ethan Fetaya · 2022
Later among the works it cites.
Cross-task knowledge distillation in multi-task recommendation
Chenxiao Yang, Junwei Pan, Xiaofeng Gao, Tingyu Jiang, Dapeng Liu, and Guihai Chen · 2022
Later among the works it cites.
A survey on negative transfer
Wen Zhang, Lingfei Deng, Lei Zhang, and Dongrui Wu · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adashare: Learning what to share for efficient deep multi-task learning
Ximeng Sun, Rameswar Panda, Rogerio Feris, and Kate Saenko · 2020
Cited alongside, same era.
Progressive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations
Hongyan Tang, Junning Liu, Ming Zhao, and Xudong Gong · 2020
Cited alongside, same era.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Cited alongside, same era.
Swad: Domain generalization by seeking flat minima
Junbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho, Seunghyun Park, Yunsung Lee, and Sungrae Park · 2021
Cited alongside, same era.
Mssm: A multiple-level sparse sharing model for efficient multi-task learning
Ke Ding, Xin Dong, Yong He, Lei Cheng, Chilin Fu, Zhaoxin Huan, Hai Li, Tan Yan, Liang Zhang, Xiaolu Zhang, and Linjian Mo · 2021
Cited alongside, same era.
Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning
Hussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy, Yihua Chen, Rahul Mazumder, Lichan Hong, and Ed Chi · 2021
Cited alongside, same era.
Multi-task distillation: Towards mitigating the negative transfer in multi-task learning
Ze Meng, Xin Yao, and Lifeng Sun · 2021
Cited alongside, same era.
Revisiting scalarization in multi-task learning: A theoretical perspective
Yuzheng Hu, Ruicheng Xian, Qilong Wu, Qiuling Fan, Lang Yin, and Han Zhao · 2023
Later among the works it cites.
Dataless knowledge fusion by merging weights of language models
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng · 2023
Later among the works it cites.
Task arithmetic in the tangent space: Improved editing of pre-trained models
Guillermo Ortiz-Jimenez, Alessandro Favero, and Pascal Frossard · 2023
Later among the works it cites.
Wenju Sun, Qingyong Li, Wen Wang, and Yangli-ao Geng · 2023
Later among the works it cites.
Hacking task confounder in meta-learning
Jingyao Wang, Yi Ren, Zeen Song, Jianqi Zhang, Changwen Zheng, and Wenwen Qiang · 2023
Later among the works it cites.
TIES-merging: Resolving interference when merging models
Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal · 2023
Later among the works it cites.
Adatask: A task-aware adaptive learning rate approach to multi-task learning
Enneng Yang, Junwei Pan, Ximei Wang, Haibin Yu, Li Shen, Xihua Chen, Lei Xiao, Jie Jiang, and Guibing Guo · 2023
Later among the works it cites.
Merging multi-task models via weight-ensembling mixture of experts
Anke Tang, Li Shen, Yong Luo, Nan Yin, Lefei Zhang, and Dacheng Tao · 2024
Later among the works it cites.
Localizing task information for improved model merging and compression
Ke Wang, Nikolaos Dimitriadis, Guillermo Ortiz-Jimenez, François Fleuret, and Pascal Frossard · 2024
Later among the works it cites.
Mixture of loRA experts
Xun Wu, Shaohan Huang, and Furu Wei · 2024
Later among the works it cites.