Fetching the paper…
Reading the bibliography…
The rapid advancement of foundation modelslarge-scale neural networks trained on diverse, extensive datasetshas revolutionized artificial intelligence, enabling unprecedented advancements across domains such as natural language processing, computer vision, and scientific discovery.
N. Kumari, B. Zhang, R. Zhang, E. Shechtman, and J.-Y. Zhu, “Multi-concept customization of text-to-image diffusion,” in CVPR , 2023, pp. 1931–1941
1941
Earlier work this paper cites.
S. Dou, E. Zhou, Y. Liu, S. Gao, W. Shen, L. Xiong, Y. Zhou, X. Wang, Z. Xi, X. Fan et al. , “LoRAMoE: Alleviating world knowledge forgetting in large language models via moe-style plugin,” in ACL , 2024, pp. 1932–1945
1945
Earlier work this paper cites.
N. Hansen and A. Ostermeier, “Adapting arbitrary normal mutation distributions in evolution strategies: The covariance matrix adaptation,” in ICEC . IEEE, 1996, pp. 312–317
1996
Earlier work this paper cites.
M. E. Wall, A. Rechtsteiner, and L. M. Rocha, “Singular value decomposition and principal component analysis,” in A practical approach to microarray data analysis . Springer, 2003, pp. 91–109
2003
Earlier work this paper cites.
I. V. Oseledets, “Tensor-train decomposition,” SISC , vol. 33, no. 5, pp. 2295–2317, 2011
2011
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” arXiv:1211.3711 , 2012
2012
Earlier work this paper cites.
D. P. Kingma, “Adam: A method for stochastic optimization,” arXiv:1412.6980 , 2014
2014
Earlier work this paper cites.
T. Bolukbasi, K.-W. Chang, J. Y. Zou, V. Saligrama, and A. T. Kalai, “Man is to computer programmer as woman is to homemaker? debiasing word embeddings,” in NeurIPS , vol. 29, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Quantized neural networks: Training neural networks with low precision weights and activations,” in JMLR , 2017, pp. 6869–6898
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” in ICML , 2019
2019
Earlier work this paper cites.
R. Banner, Y. Nahshan, and D. Soudry, “Post training 4-bit quantization of convolutional networks for rapid-deployment,” in NeurIPS , vol. 32, 2019
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in NeurIPS , vol. 33, 2020
2020
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in NeurIPS , vol. 33, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “Zero: Memory optimizations toward training trillion parameter models,” in SC20 . IEEE, 2020, pp. 1–16
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
W. Y. B. Lim, N. C. Luong, D. T. Hoang, Y. Jiao, Y.-C. Liang, Q. Yang, D. Niyato, and C. Miao, “Federated learning in mobile edge networks: A comprehensive survey,” IEEE communications surveys & tutorials , vol. 22, no. 3, pp. 2031–2063, 2020
2020
Earlier work this paper cites.
D. Alexey, “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv: 2010.11929 , 2020
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in NeurIPS , vol. 33, 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV , 2021, pp. 10 012–10 022
2021
Earlier work this paper cites.
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko et al. , “Highly accurate protein structure prediction with alphafold,” Nature , vol. 596, no. 7873, pp. 583–589, 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, “On the dangers of stochastic parrots: Can language models be too big?” Proc. 2021 ACM Conf. Fairness, Accountability, and Transparency , pp. 610–623, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Elnaggar, M. Heinzinger, C. Dallago, G. Rihawi, Y. Wang, L. Jones, T. Gibbs, T. Feher, C. Angerer, M. Steinegger, D. Bhowmik, and B. Rost, “Prottrans: Towards cracking the language of life’s code through self-supervised deep learning and high performance computing,” 2021
2021
Earlier work this paper cites.
B. Lim and S. Zohren, “Time-series forecasting with deep learning: a survey,” Philosophical Transactions of the Royal Society A , vol. 379, no. 2194, p. 20200209, 2021
2021
Earlier work this paper cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in CVPR , 2022, pp. 16 000–16 009
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” CVPR , pp. 10 684–10 695, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
H. Liu, D. Tam, M. Muqeeth, J. Mohta, T. Huang, M. Bansal, and C. A. Raffel, “Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning,” in NeurIPS , vol. 35, 2022
2022
Earlier work this paper cites.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in ICML , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” JMLR , vol. 23, no. 120, pp. 1–39, 2022
2022
Earlier work this paper cites.
S. Horvóth, C.-Y. Ho, L. Horvath, A. N. Sahu, M. Canini, and P. Richtárik, “Natural compression for distributed deep learning,” in Mathematical and Scientific Machine Learning . PMLR, 2022, pp. 129–141
2022
Earlier work this paper cites.
Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhu, W. Lu, A. dos Santos Costa, M. Fazel-Zarandi, T. Sercu, S. Candido et al. , “Language models of protein sequences at the scale of evolution enable accurate structure prediction,” BioRxiv , vol. 2022, p. 500902, 2022
2022
Earlier work this paper cites.
M. Sun, K. Zhou, X. He, Y. Wang, and X. Wang, “Gppt: Graph pre-training and prompt tuning to generalize graph neural networks,” in KDD , 2022, pp. 1717–1727
2022
Earlier work this paper cites.
Y.-L. Sung, J. Cho, and M. Bansal, “Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks,” in CVPR , 2022, pp. 5227–5237
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo et al. , “Segment anything,” in ICCV , 2023, pp. 4015–4026
2023
Earlier work this paper cites.
Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhu, W. Lu, N. Smetanin, R. Verkuil, O. Kabeli, Y. Shmueli et al. , “Evolutionary-scale prediction of atomic level protein structure with a language model,” Science , vol. 379, no. 6637, pp. 1123–1130, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. Valipour, M. Rezagholizadeh, I. Kobyzev, and A. Ghodsi, “DyLoRA: Parameter-efficient tuning of pre-trained models using dynamic search-free low-rank adaptation,” EACL , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
S. Malladi, A. Wettig, D. Yu, D. Chen, and S. Arora, “A kernel-based view of language model fine-tuning,” in ICML , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Wen and S. Chaudhuri, “Batched low-rank adaptation of foundation models,” arXiv:2312.05677 , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Chen and D. Yang, “Unlearn what you want to forget: Efficient unlearning for llms,” in EMNLP , 2023, pp. 12 041–12 052
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
Y. J. Cho, L. Liu, Z. Xu, A. Fahrezi, M. Barnes, and G. Joshi, “Heterogeneous lora for federated fine-tuning of on-device foundation models,” in NeurIPS Workshop , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
Z. Li, N. Pang, and X. Zhao, “Instruction tuning large language models for multimodal relation extraction using lora,” in WISA . Springer, 2024, pp. 364–376
2024
Closest in time.
Y. Zhang, J. Wang, L.-C. Yu, D. Xu, and X. Zhang, “Personalized lora for human-centered text understanding,” in AAAI , vol. 38, 2024, pp. 19 588–19 596
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Q. Liu, X. Wu, X. Zhao, Y. Zhu, D. Xu, F. Tian, and Y. Zheng, “Moelora: An moe-based parameter efficient fine-tuning method for multi-task medical applications,” CoRR , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
Y. Zhu, Z. Shen, Z. Zhao, S. Wang, X. Wang, X. Zhao, D. Shen, and Q. Wang, “Melo: Low-rank adaptation is better than fine-tuning for medical image diagnosis,” in ISBI . IEEE, 2024, pp. 1–5
2024
Closest in time.
2024
Closest in time.
W. Yue, J. Zhang, K. Hu, Y. Xia, J. Luo, and Z. Wang, “Surgicalsam: Efficient class promptable surgical instrument segmentation,” in AAAI , vol. 38, 2024, pp. 6890–6898
2024
Closest in time.
2024
Closest in time.
C. Kong, A. Luo, P. Bao, Y. Yu, H. Li, Z. Zheng, S. Wang, and A. C. Kot, “Moe-ffd: Mixture of experts for generalized and parameter-efficient face forgery detection,” IEEE Trans. Dependable Secure Comput. , 2024
2024
Closest in time.
2024
Closest in time.
L. Lin, H. Fan, Z. Zhang, Y. Wang, Y. Xu, and H. Ling, “Tracking meets lora: Faster training, larger model, stronger performance,” in ECCV . Springer, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
V. Shah, N. Ruiz, F. Cole, E. Lu, S. Lazebnik, Y. Li, and V. Jampani, “Ziplora: Any subject in any style by effectively merging loras,” in ECCV . Springer, 2024, pp. 422–438
2024
Closest in time.
Y. Gu, X. Wang, J. Z. Wu, Y. Shi, Y. Chen, Z. Fan, W. Xiao, R. Zhao, S. Chang, W. Wu et al. , “Mix-of-show: Decentralized low-rank adaptation for multi-concept customization of diffusion models,” in NeurIPS , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Zhu, Y. Chen, M. Ding, P. Luo, L. Wang, and J. Wang, “MoLE: Enhancing human-centric text-to-image diffusion via mixture of low-rank experts,” in NeurIPS , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Kumar and S. Chimalakonda, “Code summarization without direct access to code-towards exploring federated llms for software engineering,” in EASE , 2024, pp. 100–109
2024
Closest in time.
2024
Closest in time.
S. Zeng, D. Wang, L. Jiang, and D. Xu, “Parameter-efficient fine-tuning on large protein language models improves signal peptide prediction,” Genome research , vol. 34, 2024
2024
Closest in time.
R. Schmirler, M. Heinzinger, and B. Rost, “Fine-tuning protein language models boosts predictions across diverse tasks,” Nature Communications , vol. 15, no. 7407, 2024
2024
Closest in time.
L. Lv, Z. Lin, H. Li, Y. Liu, J. Cui, C. Yu-Chian Chen, L. Yuan, and Y. Tian, “Prollama: A protein large language model for multi-task protein language processing,” arXiv e-prints , pp. arXiv–2402, 2024
2024
Closest in time.
J. Dagdelen, A. Dunn, S. Lee, N. Walker, A. S. Rosen, G. Ceder, K. A. Persson, and A. Jain, “Structured information extraction from scientific text with large language models,” Nature communications , p. 1418, 2024
2024
Closest in time.
Z. Yang, H. Gao, D. Gao, L. Yang, L. Yang, X. Cai, W. Ning, and G. Zhang, “Mlora: Multi-domain low-rank adaptive network for ctr prediction,” in RecSys , 2024, pp. 287–297
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Zheng, W. Chao, Z. Qiu, H. Zhu, and H. Xiong, “Harnessing large language models for text-rich sequential recommendation,” in ACM WebConf , 2024, pp. 3207–3216
2024
Closest in time.
2024
Closest in time.
J. Ji, Z. Li, S. Xu, W. Hua, Y. Ge, J. Tan, and Y. Zhang, “Genrec: Large language model for generative recommendation,” in ECIR . Springer, 2024, pp. 494–502
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Wang, X. Xiang, Y. Fan, and J.-H. Xue, “Customizing 360-degree panoramas through text-to-image diffusion models,” in WACV , 2024, pp. 4933–4943
2024
Closest in time.
2024
Closest in time.
Y. Fathullah, C. Wu, E. Lakomkin, J. Jia, Y. Shangguan, K. Li, J. Guo, W. Xiong, J. Mahadeokar, O. Kalinli et al. , “Prompting large language models with speech recognition abilities,” in ICASSP . IEEE, 2024, pp. 13 351–13 355
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Dinkel, Y. Wang, Z. Yan, J. Zhang, and Y. Wang, “Ced: Consistent ensemble distillation for audio tagging,” in ICASSP . IEEE, 2024, pp. 291–295
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Jihong, Z. Xiang, and L. CHENG, “Edge computing for real-time decision making in autonomous driving: Review of challenges, solutions, and future trends.” International Journal of Advanced Computer Science & Applications , vol. 15, no. 7, 2024
2024
Closest in time.
H. Jeon, Y. Kim, and J.-J. Kim, “L4Q: Parameter efficient quantization-aware fine-tuning on large language models,” in ACL , 2025
2025
Closest in time.
Z. Zhou, Q. Zhang, H. Kumbong, and K. Olukotun, “LowRA: Accurate and efficient loRA fine-tuning of LLMs under 2 bits,” in ICML , 2025
2025
Closest in time.
Z. Zhao, T. Shen, D. Zhu, Z. Li, J. Su, X. Wang, and F. Wu, “Merging loRAs like playing LEGO: Pushing the modularity of loRA to extremes through rank-wise clustering,” in ICLR , 2025
2025
Closest in time.
Q. Huang, T. Ko, Z. Zhuang, L. Tang, and Y. Zhang, “HiRA: Parameter-efficient hadamard high-rank adaptation for large language models,” in ICLR , 2025
2025
Closest in time.
P. Albert, F. Z. Zhang, H. Saratchandran, C. Rodriguez-Opazo, A. van den Hengel, and E. Abbasnejad, “RandloRA: Full rank parameter-efficient fine-tuning of large models,” in ICLR , 2025
2025
Closest in time.
Y. Tian, B. Zhang, Z. Tu, and D. Chu, “Adapters selector: Cross-domains and multi-tasks LoRA modules integration usage method,” in COLING , 2025
2025
Closest in time.
Q. Huang, T. Ko, L. Tang, and Y. Zhang, “Comlora: A competitive learning approach for enhancing lora,” in ICLR , 2025
2025
Closest in time.
B. Yu, Z. Yang, and X. Yi, “MoKA:parameter efficiency fine-tuning via mixture of kronecker product adaption,” in COLING , 2025
2025
Closest in time.
T. Truong, C. Nguyen, H. Nguyen, M. Le, T. Le, and N. Ho, “ReploRA: Reparameterizing low-rank adaptation via the perspective of mixture of experts,” in ICML , 2025
2025
Closest in time.
C. Fan, Z. Lu, S. Liu, C. Gu, X. Qu, W. Wei, and Y. Cheng, “Make loRA great again: Boosting loRA with adaptive singular values and mixture-of-experts optimization alignment,” in ICML , 2025
2025
Closest in time.
M. Liao, W. Chen, J. Shen, S. Guo, and H. Wan, “HMoRA: Making LLMs more effective with hierarchical mixture of loRA experts,” in ICLR , 2025
2025
Closest in time.
L. Zhao, W. Zeng, S. Xiaofeng, and H. Zhou, “MoSLD: An extremely parameter-efficient mixture-of-shared LoRAs for multi-task learning,” in COLING , 2025
2025
Closest in time.
Z. Zhao, Y. Zhou, Z. Zhang, D. Zhu, T. Shen, Z. Li, J. Yang, X. Wang, J. Su, K. Kuang, Z. Wei, F. Wu, and Y. Cheng, “Each rank could be an expert: Single-ranked mixture of experts lora for multi-task learning,” 2025
2025
Closest in time.
C. Wang, Y. Lyu, Z. Sun, and L. Jing, “Continual gradient low-rank projection fine-tuning for LLMs,” in ACL , 2025
2025
Closest in time.
Y.-Y. Qian, Y.-Z. Xu, Z.-Y. Zhang, P. Zhao, and Z.-H. Zhou, “TreeloRA: Efficient continual learning via layer-wise loRAs guided by a hierarchical gradient-similarity tree,” in ICML , 2025
2025
Closest in time.
R. Singhal, K. Ponkshe, and P. Vepakomma, “FedEx-LoRA: Exact aggregation for federated and efficient fine-tuning of large language models,” in ACL , 2025
2025
Closest in time.
J. Koo, M. Jang, and J. Ok, “Towards robust and efficient federated low-rank adaptation with heterogeneous clients,” in ACL , 2025
2025
Closest in time.
Y. Zhang, H. Zhu, A. Z. Tan, D. Yu, L. Huang, and H. Yu, “pfedmxf: Personalized federated class-incremental learning with mixture of frequency aggregation,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 30 640–30 650
2025
Closest in time.
2025
Closest in time.
S. Xia, L. Sun, T. Sun, and Q. Li, “DragloRA: Online optimization of loRA adapters for drag-based image editing in diffusion model,” in ICML , 2025
2025
Closest in time.
S. Zhuang, Y. Guo, Y. Ding, K. Li, X. Chen, Y. Wang, F. Wang, Y. Zhang, C. Li, and Y. Wang, “Timestep master: Asymmetrical mixture of timestep loRA experts for versatile and efficient diffusion models in vision,” in ICML , 2025
2025
Closest in time.
C. Wang, S. Hu, G. Tan, and W. Jia, “ELoRA: Low-rank adaptation for equivariant GNNs,” in ICML , 2025
2025
Closest in time.
2025
Closest in time.
L. Pan, Z. Chen, H. Li, G. Liu, Z. Xu, Z. Liu, H. Wang, and Y. Wei, “Mixture of low rank adaptation with partial parameter sharing for time series forecasting,” arXiv , 2025
2025
Closest in time.
C. Ge, X. Wang, Z. Zhang, H. Chen, J. Fan, L. Huang, H. Xue, and W. Zhu, “Dynamic mixture of curriculum loRA experts for continual multimodal instruction tuning,” in ICML , 2025
2025
Closest in time.