Fetching the paper…
Reading the bibliography…
Recently, there has been a widespread proliferation of "expert" language models that are specialized to a specific task or domain through parameter-efficient fine-tuning.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A · 2017
Earlier work this paper cites.
Federated multi-task learning
Smith, V., Chiang, C.-K., Sanjabi, M., and Talwalkar, A. S · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Earlier work this paper cites.
Adapterfusion: Non-destructive task composition for transfer learning
Pfeiffer, J., Kamath, A., Rücklé, A., Cho, K., and Gurevych, I · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Towards a unified view of parameter-efficient transfer learning
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Earlier work this paper cites.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Earlier work this paper cites.
Fusing finetuned models for better pretraining
Choshen, L., Venezian, E., Slonim, N., and Katz, Y · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Earlier work this paper cites.
Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models
Ding, N., Qin, Y., Yang, G., Wei, F., Yang, Z., Su, Y., Hu, S., Chen, Y., Chan, C.-M., Chen, W., Yi, J., Zhao, W., Wang, X., Liu, Z., Zheng, H.-T., Chen, J., Liu, Y., Tang, J., Li, J., and Sun, M · 2022
Earlier work this paper cites.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
Sparsely activated mixture-of-experts are robust multi-task learners
Gupta, S., Mukherjee, S., Subudhi, K., Gonzalez, E., Jose, D., Awadallah, A. H., and Gao, J · 2022
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Cited alongside, same era.
Lorahub: Efficient cross-task generalization via dynamic lora composition
Huang, C., Liu, Q., Lin, B. Y., Pang, T., Du, C., and Lin, M · 2023
Later among the works it cites.
Exploring the benefits of training expert language models over instruction tuning
Jang, J., Kim, S., Ye, S., Kim, D., Logeswaran, L., Lee, M., Lee, K., and Seo, M · 2023
Later among the works it cites.
Scaling down to scale up: A guide to parameter-efficient fine-tuning
Lialin, V., Deshpande, V., and Rumshisky, A · 2023
Later among the works it cites.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2022
Cited alongside, same era.
Dataless knowledge fusion by merging weights of language models
Jin, X., Ren, X., Preotiuc-Pietro, D., and Cheng, P · 2022
Cited alongside, same era.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C · 2022
Cited alongside, same era.
Peft: State-of-the-art parameter-efficient fine-tuning methods
Mangrulkar, S., Gugger, S., Debut, L., Belkada, Y., Paul, S., and Bossan, B · 2022
Cited alongside, same era.
Merging models with fisher-weighted averaging
Matena, M. S. and Raffel, C. A · 2022
Cited alongside, same era.
Cross-task generalization via natural language crowdsourcing instructions
Mishra, S., Khashabi, D., Baral, C., and Hajishirzi, H · 2022
Cited alongside, same era.
Recycling diverse models for out-of-distribution generalization
Ramé, A., Ahuja, K., Zhang, J., Cord, M., Bottou, L., and Lopez-Paz, D · 2022
Cited alongside, same era.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., , and Wei, J · 2022
Cited alongside, same era.
Lu, K., Yuan, H., Lin, R., Lin, J., Yuan, Z., Zhou, C., and Zhou, J · 2023
Later among the works it cites.
Llama-2, mo’ lora
Maxine · 2023
Later among the works it cites.
Soft merging of experts with adaptive routing
Muqeeth, M., Liu, H., and Raffel, C · 2023
Later among the works it cites.
A case study of instruction tuning with mixture of parameter-efficient experts
Ostapenko, O., Caccia, L., Su, Z., Le Roux, N., Charlin, L., and Sordoni, A · 2023
Later among the works it cites.
Combining parameter-efficient modules for task-level generalisation
Ponti, E. M., Sordoni, A., Bengio, Y., and Reddy, S · 2023
Later among the works it cites.
Ziplora: Any subject in any style by effectively merging loras
Shah, V., Ruiz, N., Cole, F., Lu, E., Lazebnik, S., Li, Y., and Jampani, V · 2023
Later among the works it cites.
Large language model routing with benchmark datasets
Shnitzer, T., Ou, A., Silva, M., Soule, K., Sun, Y., Solomon, J., Thompson, N., and Yurochkin, M · 2023
Later among the works it cites.
Merging by matching models in task subspaces
Tam, D., Bansal, M., and Raffel, C · 2023
Later among the works it cites.
Sam-clip: Merging vision foundation models towards semantic and spatial understanding
Wang, H., Vasu, P. K. A., Faghri, F., Vemulapalli, R., Farajtabar, M., Mehta, S., Rastegari, M., Tuzel, O., and Pouransari, H · 2023
Later among the works it cites.
pi-tuning: Transferring multimodal foundation models with optimal multi-task interpolation
Wu, C., Wang, T., Ge, Y., Lu, Z., Zhou, R., Shan, Y., and Luo, P · 2023
Later among the works it cites.
TIES-merging: Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C., and Bansal, M · 2023
Later among the works it cites.
Adamerging: Adaptive model merging for multi-task learning
Yang, E., Wang, Z., Shen, L., Liu, S., Guo, G., Wang, X., and Tao, D · 2023
Later among the works it cites.
Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning
Zadouri, T., Üstün, A., Ahmadian, A., Ermiş, B., Locatelli, A., and Hooker, S · 2023
Later among the works it cites.
airoboros: Customizable implementation of the self-instruct paper
Durbin, J · 2024
Closest in time.
LlamaIndex, a data framework for your LLM applications
Liu, J · 2024
Closest in time.
Towards modular llms by building and reusing a library of loras
Ostapenko, O., Su, Z., Ponti, E. M., Charlin, L., Roux, N. L., Pereira, M., Caccia, L., and Sordoni, A · 2024
Closest in time.
Wu, X., Huang, S., and Wei, F · 2024
Closest in time.