Fetching the paper…
Reading the bibliography…
Multi-task learning (MTL) compresses the information from multiple tasks into a unified backbone to improve computational efficiency and generalization.
On the mathematical foundations of theoretical statistics
Fisher, R. A · 1922
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
The mnist database of handwritten digits
LeCun, Y · 1998
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Collobert, R. and Weston, J · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
A graphbased framework for multi-task multi-view learning
He, J. and Lawrence, R · 2011
Earlier work this paper cites.
The german traffic sign recognition benchmark: a multi-class classification competition
Stallkamp, J., Schlipsing, M., Salmen, J., and Igel, C · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval, N · 2011
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A · 2014
Earlier work this paper cites.
Multi-task policy search for robotics
Deisenroth, M. P., Englert, P., Peters, J., and Fox, D · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Multi-task learning for multiple language translation
Dong, D., Wu, H., He, W., Yu, D., and Wang, H · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Recurrent neural network for text classification with multi-task learning
Liu, P., Qiu, X., and Huang, X · 2016
Earlier work this paper cites.
Cross-stitch networks for multi-task learning
Misra, I., Shrivastava, A., Gupta, A., and Hebert, M · 2016
Earlier work this paper cites.
Sun database: Exploring a large collection of scene categories
Xiao, J., Ehinger, K. A., Hays, J., Torralba, A., and Oliva, A · 2016
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
Cheng, G., Han, J., and Lu, X · 2017
Earlier work this paper cites.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Chen, Z., Badrinarayanan, V., Lee, C.-Y., and Rabinovich, A · 2018
Earlier work this paper cites.
Rank and rate: multi-task learning for recommender systems
Hadash, G., Shalom, O. S., and Osadchy, R · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Kendall, A., Gal, Y., and Cipolla, R · 2018
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Ma, J., Zhao, Z., Yi, X., Chen, J., Hong, L., and Chi, E. H · 2018
Earlier work this paper cites.
Multi-task learning as multi-objective optimization
Sener, O. and Koltun, V · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Helber, P., Bischke, B., Dengel, A., and Borth, D · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Cited alongside, same era.
End-to-end multi-task learning with attention
Liu, S., Johns, E., and Davison, A. J · 2019
Cited alongside, same era.
Snr: Sub-network routing for flexible parameter sharing in multi-task learning
Ma, J., Zhao, Z., Chen, J., Li, A., Hong, L., and Chi, E. H · 2019
Cited alongside, same era.
Predicting different types of conversions with multi-task learning in online advertising
Pan, J., Mao, Y., Ruiz, A. L., Sun, Y., and Flores, A · 2019
Cited alongside, same era.
Just pick a sign: Optimizing deep multitask models with gradient sign dropout
Merging models with fisher-weighted averaging
Matena, M. S. and Raffel, C. A · 2022
Later among the works it cites.
Multi-task learning as a bargaining game
Navon, A., Shamsian, A., Achituve, I., Maron, H., Kawaguchi, K., Chechik, G., and Fetaya, E · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al · 2022
Later among the works it cites.
A survey on negative transfer
Zhang, W., Deng, L., Zhang, L., and Wu, D · 2022
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
Ainsworth, S., Hayase, J., and Srinivasa, S · 2023
Later among the works it cites.
Revisiting scalarization in multi-task learning: A theoretical perspective
Hu, Y., Xian, R., Wu, Q., Fan, Q., Yin, L., and Zhao, H · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, Z., Ngiam, J., Huang, Y., Luong, T., Kretzschmar, H., Chai, Y., and Anguelov, D · 2020
Cited alongside, same era.
Mtl-nas: Task-agnostic neural architecture search towards general-purpose multi-task learning
Gao, Y., Bai, H., Jie, Z., Ma, J., Jia, K., and Liu, W · 2020
Cited alongside, same era.
Stochastic weight averaging in parallel: Large-batch training that generalizes well
Gupta, V., Serrano, S. A., and DeCoste, D · 2020
Cited alongside, same era.
Knowledge distillation for multi-task learning
Li, W.-H. and Bilen, H · 2020
Cited alongside, same era.
Adashare: Learning what to share for efficient deep multi-task learning
Sun, X., Panda, R., Feris, R., and Saenko, K · 2020
Cited alongside, same era.
Progressive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations
Tang, H., Liu, J., Zhao, M., and Gong, X · 2020
Cited alongside, same era.
Gradient surgery for multi-task learning
Yu, T., Kumar, S., Gupta, A., Levine, S., Hausman, K., and Finn, C · 2020
Cited alongside, same era.
Later among the works it cites.
Lorahub: Efficient cross-task generalization via dynamic lora composition
Huang, C., Liu, Q., Lin, B. Y., Pang, T., Du, C., and Lin, M · 2023
Later among the works it cites.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2023
Later among the works it cites.
Dataless knowledge fusion by merging weights of language models
Jin, X., Ren, X., Preotiuc-Pietro, D., and Cheng, P · 2023
Later among the works it cites.
Task arithmetic in the tangent space: Improved editing of pre-trained models
Ortiz-Jimenez, G., Favero, A., and Frossard, P · 2023
Later among the works it cites.
Re-basin via implicit sinkhorn differentiation
Peña, F. A. G., Medeiros, H. R., Dubail, T., Aminbeidokhti, M., Granger, E., and Pedersoli, M · 2023
Later among the works it cites.
Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Ramé, A., Couairon, G., Shukor, M., Dancette, C., Gaya, J.-B., Soulier, L., and Cord, M · 2023
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2023
Later among the works it cites.
Zipit! merging models from different tasks without training
Stoica, G., Bolya, D., Bjorner, J., Hearn, T., and Hoffman, J · 2023
Later among the works it cites.
Concrete subspace learning based interference elimination for multi-task model fusion
Tang, A., Shen, L., Luo, Y., Ding, L., Hu, H., Du, B., and Tao, D · 2023
Later among the works it cites.
Multi-task deep recommender systems: A survey
Wang, Y., Lam, H. T., Wong, Y., Liu, Z., Zhao, X., Wang, Y., Chen, B., Guo, H., and Tang, R · 2023
Later among the works it cites.
pi-tuning: Transferring multimodal foundation models with optimal multi-task interpolation
Wu, C., Wang, T., Ge, Y., Lu, Z., Zhou, R., Shan, Y., and Luo, P · 2023
Later among the works it cites.
Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C., and Bansal, M · 2023
Later among the works it cites.
Adatask: A task-aware adaptive learning rate approach to multi-task learning
Yang, E., Pan, J., Wang, X., Yu, H., Shen, L., Chen, X., Xiao, L., Jiang, J., and Guo, G · 2023
Later among the works it cites.
Language models are super mario: Absorbing abilities from homologous models as a free lunch
Yu, L., Yu, B., Yu, H., Huang, F., and Li, Y · 2023
Later among the works it cites.
Composing parameter-efficient modules with arithmetic operations
Zhang, J., Chen, S., Liu, J., and He, J · 2023
Later among the works it cites.
Multi-scenario and multi-task aware feature interaction for recommendation system
Song, D., Yang, E., Guo, G., Shen, L., Jiang, L., and Wang, X · 2024
Closest in time.
Parameter efficient multi-task model fusion with partial linearization
Tang, A., Shen, L., Luo, Y., Zhan, Y., Hu, H., Du, B., Chen, Y., and Tao, D · 2024
Closest in time.
Adamerging: Adaptive model merging for multi-task learning
Yang, E., Wang, Z., Shen, L., Liu, S., Guo, G., Wang, X., and Tao, D · 2024
Closest in time.