Fetching the paper…
Reading the bibliography…
Fine-tuning large language models (LMs) for individual tasks yields strong performance but is expensive for deployment and storage.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu. 2019 · 1907
Earlier work this paper cites.
Procrustes problems , volume 30
John C Gower and Garmt B Dijksterhuis. 2004 · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. 2007 · 2007
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
First quora dataset release: Question pairs
Shankar Iyer, Nikhil Dandekar, Kornél Csernai, et al. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I Loshchilov. 2017 · 2017
Earlier work this paper cites.
Multi-task learning as multi-objective optimization
Ozan Sener and Vladlen Koltun. 2018 · 2018
Earlier work this paper cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
A survey on multi-task learning
Yu Zhang and Qiang Yang. 2021 · 2021
Cited alongside, same era.
Git re-basin: Merging models modulo permutation symmetries
Samuel K Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa. 2022 · 2022
Cited alongside, same era.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2022 · 2022
Ethos: Rectifying language models in orthogonal parameter space
Lei Gao, Yue Niu, Tingting Tang, Salman Avestimehr, and Murali Annavaram. 2024 · 2024
Later among the works it cites.
Localize-and-stitch: Efficient model merging via sparse task arithmetic
Yifei He, Yuzheng Hu, Yong Lin, Tong Zhang, and Han Zhao. 2024 · 2024
Later among the works it cites.
Emr-merging: Tuning-free high-performance model merging
Chenyu Huang, Peng Ye, Tao Chen, Tong He, Xiangyu Yue, and Wanli Ouyang. 2024 · 2024
Later among the works it cites.
Agent skill acquisition for large language models via cycleqd
So Kuroki, Taishi Nakamura, Takuya Akiba, and Yujin Tang. 2024 · 2024
Later among the works it cites.
Twin-merging: Dynamic integration of modular expertise in model merging
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dataless knowledge fusion by merging weights of language models
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng. 2022 · 2022
Cited alongside, same era.
Merging models with fisher-weighted averaging
Michael S Matena and Colin A Raffel. 2022 · 2022
Cited alongside, same era.
Byom: Building your own multi-task model for free
Weisen Jiang, Baijiong Lin, Han Shi, Yu Zhang, Zhenguo Li, and James Kwok. 2023 · 2023
Cited alongside, same era.
Task arithmetic in the tangent space: Improved editing of pre-trained models
Guillermo Ortiz-Jimenez, Alessandro Favero, and Pascal Frossard. 2023 · 2023
Cited alongside, same era.
Zipit! merging models from different tasks without training
George Stoica, Daniel Bolya, Jakob Bjorner, Pratik Ramesh, Taylor Hearn, and Judy Hoffman. 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Cited alongside, same era.
Model breadcrumbs: Scaling multi-task model merging with sparse masks
MohammadReza Davari and Eugene Belilovsky. 2024 · 2024
Cited alongside, same era.
Zhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu, Dangyang Chen, and Yu Cheng. 2024 · 2024
Later among the works it cites.
Lora soups: Merging loras for practical skill composition tasks
Akshara Prabhakar, Yuanzhi Li, Karthik Narasimhan, Sham Kakade, Eran Malach, and Samy Jelassi. 2024 · 2024
Later among the works it cites.
Model merging with svd to tie the knots
George Stoica, Pratik Ramesh, Boglarka Ecsedi, Leshem Choshen, and Judy Hoffman. 2024 · 2024
Later among the works it cites.
Merging by matching models in task parameter subspaces
Derek Tam, Mohit Bansal, and Colin Raffel. 2024 · 2024
Later among the works it cites.
Task arithmetic through the lens of one-shot federated learning
Zhixu Tao, Ian Mason, Sanjeev Kulkarni, and Xavier Boix. 2024 · 2024
Later among the works it cites.
Ties-merging: Resolving interference when merging models
Prateek Yadav, Derek Tam, Leshem Choshen, Colin A Raffel, and Mohit Bansal. 2024 · 2024
Later among the works it cites.
Ziyu Zhao, Tao Shen, Didi Zhu, Zexi Li, Jing Su, Xuwu Wang, Kun Kuang, and Fei Wu. 2024 · 2024
Later among the works it cites.
Metagpt: Merging large language models using model exclusive task arithmetic
Yuyan Zhou, Liang Song, Bingning Wang, and Weipeng Chen. 2024 · 2024
Later among the works it cites.
Smarter fine-tuning: How lora enhances large language models
Yue Gang, Jianhong Shun, and Mu Qing. 2025 · 2025
Closest in time.