Fetching the paper…
Reading the bibliography…
The growing number of parameter-efficient adaptations of a base large language model (LLM) calls for studying whether we can reuse such trained adapters to improve performance for new tasks.
Bagging predictors
Breiman, L · 1996
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
Ensemble methods in machine learning
Dietterich, T. G · 2000
Earlier work this paper cites.
Automatic evaluation of summaries using n-gram co-occurrence statistics
Lin, C.-Y. and Hovy, E · 2003
Earlier work this paper cites.
Exploring and predicting transferability across nlp tasks
Vu, T., Wang, T., Munkhdalai, T., Sordoni, A., Trischler, A., Mattarella-Micke, A., Maji, S., and Iyyer, M · 2005
Earlier work this paper cites.
Hard mixtures of experts for large scale weakly supervised vision
Gross, S., Ranzato, M., and Szlam, A · 2017
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
Caron, M., Bojanowski, P., Joulin, A., and Douze, M · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Nevergrad - A gradient-free optimization platform
Rapin, J. and Teytaud, O · 2018
Earlier work this paper cites.
Adaptive knowledge sharing in multi-task learning: Improving low-resource neural machine translation
Zaremoodi, P., Buntine, W., and Haffari, G · 2018
Earlier work this paper cites.
A modulation module for multi-task learning with applications in image retrieval
Zhao, X., Li, H., Shen, X., Liang, X., and Wu, Y · 2018
Earlier work this paper cites.
Stochastic filter groups for multi-task cnns: Learning specialist and generalist convolution kernels
Bragman, F. J., Tanno, R., Ourselin, S., Alexander, D. C., and Cardoso, J · 2019
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
Specializing word embeddings (for parsing) by information bottleneck
Li, X. L. and Eisner, J · 2019
Earlier work this paper cites.
The low-rank eigenvalue problem
Nakatsukasa, Y · 2019
Earlier work this paper cites.
Many task learning with task routing
Strezoski, G., Noord, N. v., and Worring, M · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Earlier work this paper cites.
Linear mode connectivity and the lottery ticket hypothesis
Frankle, J., Dziugaite, G. K., Roy, D., and Carbin, M · 2020
Earlier work this paper cites.
Privacy in deep learning: A survey
Mireshghallah, F., Taram, M., Vepakomma, P., Singh, A., Raskar, R., and Esmaeilzadeh, H · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Exploring and predicting transferability across NLP tasks
Vu, T., Wang, T., Munkhdalai, T., Sordoni, A., Trischler, A., Mattarella-Micke, A., Maji, S., and Iyyer, M · 2020
Earlier work this paper cites.
Batchensemble: An alternative approach to efficient ensemble and lifelong learning, 2020
Wen, Y., Tran, D., and Ba, J · 2020
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Earlier work this paper cites.
Adapterhub playground: Simple and flexible few-shot learning with adapters
Beck, T., Bohlender, B., Viehmann, C., Hane, V., Adamson, Y., Khuri, J., Brossmann, J., Pfeiffer, J., and Gurevych, I · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Efficient hierarchical domain adaptation for pretrained language models
Chronopoulou, A., Peters, M. E., and Dodge, J · 2021
Earlier work this paper cites.
Enslm: Ensemble language model for data diversity by semantic clustering
Duan, Z., Zhang, H., Wang, C., Wang, Z., Chen, B., and Zhou, M · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., et al · 2021
Earlier work this paper cites.
Efficiently identifying task groupings for multi-task learning
Fifty, C., Amid, E., Zhao, Z., Yu, T., Anil, R., and Finn, C · 2021
Cited alongside, same era.
Parameter-efficient multi-task fine-tuning for Transformers via shared hypernetworks
Karimi Mahabadi, R., Ruder, S., Dehghani, M., and Henderson, J · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning, 2021
Lester, B., Al-Rfou, R., and Constant, N · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Cited alongside, same era.
Continual learning via local module composition
Ostapenko, O., Rodriguez, P., Caccia, M., and Charlin, L · 2021
Cited alongside, same era.
AdapterFusion: Non-destructive task composition for transfer learning
Mixture of cluster-conditional lora experts for vision-language instruction tuning
Gou, Y., Liu, Z., Chen, K., Hong, L., Xu, H., Li, A., Yeung, D.-Y., Kwok, J. T., and Zhang, Y · 2023
Later among the works it cites.
Scaling expert language models with unsupervised domain discovery
Gururangan, S., Li, M., Lewis, M., Shi, W., Althoff, T., Smith, N. A., and Zettlemoyer, L · 2023
Later among the works it cites.
Lorahub: Efficient cross-task generalization via dynamic lora composition
Huang, C., Liu, Q., Lin, B. Y., Pang, T., Du, C., and Lin, M · 2023
Later among the works it cites.
Exploring the benefits of training expert language models over instruction tuning, 2023
Jang, J., Kim, S., Ye, S., Kim, D., Logeswaran, L., Lee, M., Lee, K., and Seo, M · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pfeiffer, J., Kamath, A., Rücklé, A., Cho, K., and Gurevych, I · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
Spot: Better frozen model adaptation through soft prompt transfer
Vu, T., Lester, B., Constant, N., Al-Rfou, R., and Cer, D · 2021
Cited alongside, same era.
Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models
Wang, Z., Tsvetkov, Y., Firat, O., and Cao, Y · 2021
Cited alongside, same era.
Git re-basin: Merging models modulo permutation symmetries
Ainsworth, S. K., Hayase, J., and Srinivasa, S · 2022
Cited alongside, same era.
Mod-squad: Designing mixture of experts as modular multi-task learners, 2022
Chen, Z., Shen, Y., Ding, M., Chen, Z., Zhao, H., Learned-Miller, E., and Gan, C · 2022
Cited alongside, same era.
Memory efficient continual learning with transformers
Ermis, B., Zappella, G., Wistuba, M., Rawal, A., and Archambeau, C · 2022
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Later among the works it cites.
Population parameter averaging (papa), 2023
Jolicoeur-Martineau, A., Gervais, E., Fatras, K., Zhang, Y., and Lacoste-Julien, S · 2023
Later among the works it cites.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al · 2023
Later among the works it cites.
Routing to the expert: Efficient reward-guided ensemble of large language models
Lu, K., Yuan, H., Lin, R., Lin, J., Yuan, Z., Zhou, C., and Zhou, J · 2023
Later among the works it cites.
Phi-2: The Surprising Power of Small Language Models, 2023
Microsoft Research · 2023
Later among the works it cites.
Soft merging of experts with adaptive routing
Muqeeth, M., Liu, H., and Raffel, C · 2023
Later among the works it cites.
A case study of instruction tuning with mixture of parameter-efficient experts
Ostapenko, O., Caccia, L., Su, Z., Le Roux, N., Charlin, L., and Sordoni, A · 2023
Later among the works it cites.
Pfeiffer, J., Ruder, S., Vulić, I., and Ponti, E. M · 2023
Later among the works it cites.
Combining parameter-efficient modules for task-level generalisation
Ponti, E. M., Sordoni, A., Bengio, Y., and Reddy, S · 2023
Later among the works it cites.
Adapters: A unified library for parameter-efficient and modular transfer learning
Poth, C., Sterz, H., Paul, I., Purkayastha, S., Engländer, L., Imhof, T., Vulić, I., Ruder, S., Gurevych, I., and Pfeiffer, J · 2023
Later among the works it cites.
Model ratatouille: Recycling diverse models for out-of-distribution generalization
Ramé, A., Ahuja, K., Zhang, J., Cord, M., Bottou, L., and Lopez-Paz, D · 2023
Later among the works it cites.
Ziplora: Any subject in any style by effectively merging loras
Shah, V., Ruiz, N., Cole, F., Lu, E., Lazebnik, S., Li, Y., and Jampani, V · 2023
Later among the works it cites.
Large language model routing with benchmark datasets
Shnitzer, T., Ou, A., Silva, M., Soule, K., Sun, Y., Solomon, J., Thompson, N., and Yurochkin, M · 2023
Later among the works it cites.
Merging by matching models in task subspaces
Tam, D., Bansal, M., and Raffel, C · 2023
Later among the works it cites.
How far can camels go? exploring the state of instruction tuning on open resources
Wang, Y., Ivison, H., Dasigi, P., Hessel, J., Khot, T., Chandu, K. R., Wadden, D., MacMillan, K., Smith, N. A., Beltagy, I., et al · 2023
Later among the works it cites.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., and Goldstein, T · 2023
Later among the works it cites.
π \pi -tuning: Transferring multimodal foundation models with optimal multi-task interpolation
Wu, C., Wang, T., Ge, Y., Lu, Z., Zhou, R., Shan, Y., and Luo, P · 2023
Later among the works it cites.
Adamerging: Adaptive model merging for multi-task learning
Yang, E., Wang, Z., Shen, L., Liu, S., Guo, G., Wang, X., and Tao, D · 2023
Later among the works it cites.
Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning
Zadouri, T., Üstün, A., Ahmadian, A., Ermiş, B., Locatelli, A., and Hooker, S · 2023
Later among the works it cites.
Model merging by uncertainty-based gradient matching
Daheim, N., Möllenhoff, T., Ponti, E., Gurevych, I., and Khan, M. E · 2024
Closest in time.
Multiple-choice normalization
EleutherAI · 2024
Closest in time.
Lorahub: Efficient cross-task generalization via dynamic lora composition, 2024
Huang, C., Liu, Q., Lin, B. Y., Pang, T., Du, C., and Lin, M · 2024
Closest in time.
Mixtral of experts, 2024
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Closest in time.
Learning to route among specialized experts for zero-shot generalization
Muqeeth, M., Liu, H., Liu, Y., and Raffel, C · 2024
Closest in time.
Mole: Mixture of lora experts
Xun Wu, Shaohan Huang, F. W · 2024
Closest in time.
Ties-merging: Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C. A., and Bansal, M · 2024
Closest in time.