Fetching the paper…
Reading the bibliography…
Model merging, which combines multiple models into a single model, has gained popularity in recent years.
The MNIST database of handwritten digits, 1998
LeCun, Y · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Nonlinear Multiobjective Optimization , volume 12
Miettinen, K · 1999
Earlier work this paper cites.
Convex Optimization
Boyd, S. P. and Vandenberghe, L · 2004
Earlier work this paper cites.
Semi-supervised learning by entropy minimization
Grandvalet, Y. and Bengio, Y · 2004
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., Ng, A. Y., et al · 2011
Earlier work this paper cites.
The German traffic sign recognition benchmark: A multi-class classification competition
Stallkamp, J., Schlipsing, M., Salmen, J., and Igel, C · 2011
Earlier work this paper cites.
Multiple-gradient descent algorithm (MGDA) for multiobjective optimization
Désidéri, J.-A · 2012
Earlier work this paper cites.
3D object representations for fine-grained categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Sun database: Exploring a large collection of scene categories
Xiao, J., Ehinger, K. A., Hays, J., Torralba, A., and Oliva, A · 2016
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
Cheng, G., Han, J., and Lu, X · 2017
Earlier work this paper cites.
Multi-task learning as multi-objective optimization
Sener, O. and Koltun, V · 2018
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Helber, P., Bischke, B., Dengel, A., and Borth, D · 2019
Cited alongside, same era.
You only train once: Loss-conditional training of deep networks
Dosovitskiy, A. and Djolonga, J · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Controllable Pareto multi-task learning
Lin, X., Yang, Z., Zhang, Q., and Kwong, S · 2020
Cited alongside, same era.
Multi-task learning with user preferences: Gradient descent with controlled ascent in Pareto optimization
Mahapatra, D. and Rajan, V · 2020
Cited alongside, same era.
Git Re-Basin: Merging models modulo permutation symmetries
Ainsworth, S., Hayase, J., and Srinivasa, S · 2023
Later among the works it cites.
Pareto manifold learning: Tackling multiple tasks via ensembles of single-task models
Dimitriadis, N., Frossard, P., and Fleuret, F · 2023
Later among the works it cites.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2023
Later among the works it cites.
Dataless knowledge fusion by merging weights of language models
Jin, X., Ren, X., Preotiuc-Pietro, D., and Cheng, P · 2023
Later among the works it cites.
REPAIR: Renormalizing permuted activations for interpolation repair
Jordan, K., Sedghi, H., Saukh, O., Entezari, R., and Neyshabur, B · 2023
Later among the works it cites.
TIES-Merging: Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C. A., and Bansal, M · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Conflict-averse gradient descent for multi-task learning
Liu, B., Liu, X., Jin, X., Stone, P., and Liu, Q · 2021
Cited alongside, same era.
Learning the Pareto front with hypernetworks
Navon, A., Shamsian, A., Fetaya, E., and Chechik, G · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Scalable Pareto front approximation for deep multi-objective learning
Ruchte, M. and Grabocka, J · 2021
Cited alongside, same era.
Parameter-efficient fine-tuning for vision transformers
He, X., Li, C., Zhang, P., Yang, J., and Wang, X. E · 2022
Cited alongside, same era.
Merging models with Fisher-weighted averaging
Matena, M. S. and Raffel, C. A · 2022
Cited alongside, same era.
Later among the works it cites.
A survey on mixture of experts
Cai, W., Jiang, J., Wang, F., Tang, J., Kim, S., and Huang, J · 2024
Closest in time.
Efficient Pareto manifold learning with low-rank structure
Chen, W. and Kwok, J · 2024
Closest in time.
Model merging by uncertainty-based gradient matching
Daheim, N., Möllenhoff, T., Ponti, E., Gurevych, I., and Khan, M. E · 2024
Closest in time.
QLoRA: Efficient finetuning of quantized LLMs
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2024
Closest in time.
EMR-merging: Tuning-free high-performance model merging
Huang, C., Ye, P., Chen, T., He, T., Yue, X., and Ouyang, W · 2024
Closest in time.
Smooth Tchebycheff scalarization for multi-objective optimization
Lin, X., Zhang, X., Yang, Z., Liu, F., Wang, Z., and Zhang, Q · 2024
Closest in time.
DoRA: Weight-decomposed low-rank adaptation
Liu, S.-Y., Wang, C.-Y., Yin, H., Molchanov, P., Wang, Y.-C. F., Cheng, K.-T., and Chen, M.-H · 2024
Closest in time.
Twin-merging: Dynamic integration of modular expertise in model merging
Lu, Z., Fan, C., Wei, W., Qu, X., Chen, D., and Cheng, Y · 2024
Closest in time.
Rewarded soups: Towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Rame, A., Couairon, G., Dancette, C., Gaya, J.-B., Shukor, M., Soulier, L., and Cord, M · 2024
Closest in time.
ZipIt! Merging models from different tasks without training
Stoica, G., Bolya, D., Bjorner, J. B., Ramesh, P., Hearn, T., and Hoffman, J · 2024
Closest in time.