Fetching the paper…
Reading the bibliography…
Recently, various merging methods have been proposed to build a multi-task model from task-specific finetuned models without retraining.
Unsupervised construction of large paraphrase corpora: exploiting massively parallel news sources
Dolan, B., Quirk, C., and Brockett, C · 2004
Earlier work this paper cites.
Visualizing data using t-SNE
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
The MNIST handwritten digit database
LeCun, Y., Cortes, C., and Burges, C · 2010
Earlier work this paper cites.
Torchvision the machine-vision package of Torch
Marcel, S. and Rodriguez, Y · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y · 2011
Earlier work this paper cites.
The German traffic sign recognition benchmark: a multi-class classification competition
Stallkamp, J., Schlipsing, M., Salmen, J., and Igel, C · 2011
Earlier work this paper cites.
3D object representations for fine-grained categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A · 2014
Earlier work this paper cites.
Multi-task learning for multiple language translation
Dong, D., Wu, H., He, W., Yu, D., and Wang, H · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Exploring sparsity in recurrent neural networks
Narang, S., Diamos, G., Sengupta, S., and Elsen, E · 2016
Earlier work this paper cites.
SUN database: Exploring a large collection of scene categories
Xiao, J., Ehinger, K. A., Hays, J., Torralba, A., and Oliva, A · 2016
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
Cheng, G., Han, J., and Lu, X · 2017
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Kendall, A., Gal, Y., and Cipolla, R · 2018
Earlier work this paper cites.
Rethinking the value of network pruning
Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T · 2018
Earlier work this paper cites.
MODNet: Motion and appearance based moving object detection network for autonomous driving
Siam, M., Mahgoub, H., Zahran, M., Yogamani, S., Jagersand, M., and El-Sallab, A · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Cited alongside, same era.
EuroSAT: A novel dataset and deep learning benchmark for land use and land cover classification
Helber, P., Bischke, B., Dengel, A., and Borth, D · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Cited alongside, same era.
End-to-end multi-task learning with attention
Liu, S., Johns, E., and Davison, A. J · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
PyTorch image models
Wightman, R · 2019
Reasonable effectiveness of random weighting: A litmus test for multi-task learning
Lin, B., Ye, F., Zhang, Y., and Tsang, I · 2022
Later among the works it cites.
Merging models with fisher-weighted averaging
Matena, M. S. and Raffel, C. A · 2022
Later among the works it cites.
MetaICL: Learning to learn in context
Min, S., Lewis, M., Zettlemoyer, L., and Hajishirzi, H · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
Structured pruning learns compact and accurate models
Xia, M., Zhong, Z., and Chen, D · 2022
Later among the works it cites.
A survey on multi-task learning
Zhang, Y. and Yang, Q · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A large-scale study of representation learning with the visual task adaptation benchmark
Zhai, X., Puigcerver, J., Kolesnikov, A., Ruyssen, P., Riquelme, C., Lucic, M., Djolonga, J., Pinto, A. S., Neumann, M., Dosovitskiy, A., Beyer, L., Bachem, O., Tschannen, M., Michalski, M., Bousquet, O., Gelly, S., and Houlsby, N · 2019
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Pruning from scratch
Wang, Y., Zhang, X., Xie, L., Zhou, J., Su, H., Zhang, B., and Hu, X · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
Advancing model pruning via bi-level optimization
Zhang, Y., Yao, Y., Ram, P., Zhao, P., Chen, T., Hong, M., Wang, Y., and Liu, S · 2022
Later among the works it cites.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2023
Closest in time.
Effective structured-prompting by meta-learning and representitive verbalizer
Jiang, W., Zhang, Y., and Kwok, J · 2023
Closest in time.
Dataless knowledge fusion by merging weights of language models
Jin, X., Ren, X., Preotiuc-Pietro, D., and Cheng, P · 2023
Closest in time.
Dual-balancing for multi-task learning
Lin, B., Jiang, W., Ye, F., Zhang, Y., Chen, P., Chen, Y.-C., Liu, S., and Kwok, J. T · 2023
Closest in time.
DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K · 2023
Closest in time.
Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C., and Bansal, M · 2023
Closest in time.
MetaMath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., Yu, J., Liu, Z., Zhang, Y., Kwok, J. T., Li, Z., Weller, A., and Liu, W · 2023
Closest in time.
Scaling relationship on learning mathematical reasoning with large language models
Yuan, Z., Yuan, H., Li, C., Dong, G., Tan, C., and Zhou, C · 2023
Closest in time.
Model merging by uncertainty-based gradient matching
Daheim, N., Möllenhoff, T., Ponti, E. M., Gurevych, I., and Khan, M. E · 2024
Closest in time.
Adamerging: Adaptive model merging for multi-task learning
Yang, E., Wang, Z., Shen, L., Liu, S., Guo, G., Wang, X., and Tao, D · 2024
Closest in time.
A first-order multi-gradient algorithm for multi-objective bi-level optimization
Ye, F., Lin, B., Cao, X., Zhang, Y., and Tsang, I · 2024
Closest in time.