Fetching the paper…
Reading the bibliography…
Parameter-efficient fine-tuning stands as the standard for efficiently fine-tuning large language and vision models on downstream tasks.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Xiao, J., Hays, J., Ehinger, K. A., Oliva, A., and Torralba, A · 2010
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Bossard, L., Guillaumin, M., and Van Gool, L · 2014
Earlier work this paper cites.
Deeper, broader and artier domain generalization
Li, D., Yang, Y., Song, Y.-Z., and Hospedales, T. M · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Rebuffi, S.-A., Bilen, H., and Vedaldi, A · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Welbl, J., Liu, N. F., and Gardner, M · 2017
Earlier work this paper cites.
Simple, scalable adaptation for neural machine translation
Bapna, A., Arivazhagan, N., and Firat, O · 2019
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Helber, P., Bischke, B., Dengel, A., and Borth, D · 2019
Earlier work this paper cites.
Do better imagenet models transfer better?
Kornblith, S., Shlens, J., and Le, Q. V · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
Parameter-efficient transfer learning with diff pruning
Guo, D., Rush, A. M., and Kim, Y · 2020
Earlier work this paper cites.
Exploring versatile generative language model via parameter-efficient transfer learning
Lin, Z., Madotto, A., and Fung, P · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Anatomy of catastrophic forgetting: Hidden representations and task semantics
Ramasesh, V. V., Dyer, E., and Raghu, M · 2020
Cited alongside, same era.
Getting closer to ai complete question answering: A set of prerequisite real tasks
Rogers, A., Kovaleva, O., Downey, M., and Rumshisky, A · 2020
Cited alongside, same era.
Bhakthavatsalam, S., Khashabi, D., Khot, T., Mishra, B. D., Richardson, K., Sabharwal, A., Schoenick, C., Tafjord, O., and Clark, P · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Cited alongside, same era.
Probing representation forgetting in supervised and unsupervised continual learning
Davari, M., Asadi, N., Mudur, S., Aljundi, R., and Belilovsky, E · 2022
Later among the works it cites.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2022
Later among the works it cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C. A · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al · 2022
Later among the works it cites.
Multi-head adapter routing for cross-task generalization
Caccia, L., Ponti, E., Su, Z., Pereira, M., Le Roux, N., and Sordoni, A · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Cited alongside, same era.
Merging models with fisher-weighted averaging, 2021
Matena, M. and Raffel, C · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Imagenet-21k pretraining for the masses
Ridnik, T., Ben-Baruch, E., Noy, A., and Zelnik-Manor, L · 2021
Cited alongside, same era.
Training neural networks with fixed sparse masks
Sung, Y.-L., Nair, V., and Raffel, C. A · 2021
Cited alongside, same era.
Crossfit: A few-shot learning challenge for cross-task generalization in nlp
Ye, Q., Lin, B. Y., and Ren, X · 2021
Cited alongside, same era.
Later among the works it cites.
Model breadcrumbs: Scaling multi-task model merging with sparse masks
Davari, M. and Belilovsky, E · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Later among the works it cites.
Concept sliders: Lora adaptors for precise control in diffusion models
Gandikota, R., Materzynska, J., Zhou, T., Torralba, A., and Bau, D · 2023
Later among the works it cites.
Lorahub: Efficient cross-task generalization via dynamic lora composition
Huang, C., Liu, Q., Lin, B. Y., Pang, T., Du, C., and Lin, M · 2023
Later among the works it cites.
Combining parameter-efficient modules for task-level generalisation
Ponti, E. M., Sordoni, A., Bengio, Y., and Reddy, S · 2023
Later among the works it cites.
Model ratatouille: Recycling diverse models for out-of-distribution generalization
Ramé, A., Ahuja, K., Zhang, J., Cord, M., Bottou, L., and Lopez-Paz, D · 2023
Later among the works it cites.
Ziplora: Any subject in any style by effectively merging loras
Shah, V., Ruiz, N., Cole, F., Lu, E., Lazebnik, S., Li, Y., and Jampani, V · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Baize: An open-source chat model with parameter-efficient tuning on self-chat data
Xu, C., Guo, D., Duan, N., and McAuley, J · 2023
Later among the works it cites.