Fetching the paper…
Reading the bibliography…
Low-Rank Adaptation (LoRA) emerges as a popular parameter-efficient fine-tuning (PEFT) method, which proposes to freeze pretrained model weights and update an additive low-rank trainable matrix.
An introduction to hyperplane arrangements
Stanley, R. P. et al · 2004
Earlier work this paper cites.
Robust principal component analysis?, 2009
Candes, E. J., Li, X., Ma, Y., and Wright, J · 2009
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Recht, B., Fazel, M., and Parrilo, P. A · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
A riemannian geometry for low-rank matrix completion, 2012
Mishra, B., Apuroop, K. A., and Sepulchre, R · 2012
Earlier work this paper cites.
Riemannian preconditioning
Mishra, B. and Sepulchre, R · 2016
Earlier work this paper cites.
The WebNLG challenge: Generating text from RDF data
Gardent, C., Shimorina, A., Narayan, S., and Perez-Beltrachini, L · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Kingma, D. P. and Ba, J · 2017
Earlier work this paper cites.
The e2e dataset: New challenges for end-to-end generation, 2017
Novikova, J., Dušek, O., and Rieser, V · 2017
Earlier work this paper cites.
Deep information propagation, 2017
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J · 2017
Earlier work this paper cites.
Shampoo: Preconditioned stochastic tensor optimization, 2018
Gupta, V., Koren, T., and Singer, Y · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Earlier work this paper cites.
On the impact of the activation function on deep neural networks training, 2019
Hayou, S., Doucet, A., and Rousseau, J · 2019
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Neural networks are convex regularizers: Exact polynomial-time convex optimization formulations for two-layer networks, 2020
Pilanci, M. and Ergen, T · 2020
Cited alongside, same era.
Scaling limits of wide neural networks with weight sharing: Gaussian process behavior, gradient independence, and neural tangent kernel derivation, 2020
Yang, G · 2020
Cited alongside, same era.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
DART: Open-domain structured data record to text generation
Qlora: Efficient finetuning of quantized llms, 2023
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Later among the works it cites.
Mix-of-show: Decentralized low-rank adaptation for multi-concept customization of diffusion models, 2023
Gu, Y., Wang, X., Wu, J. Z., Shi, Y., Chen, Y., Fan, Z., Xiao, W., Zhao, R., Chang, S., Wu, W., Ge, Y., Shan, Y., and Shou, M. Z · 2023
Later among the works it cites.
Preconditioning matters: Fast global convergence of non-convex matrix factorization via scaled gradient descent
Jia, X., Wang, H., Peng, J., Feng, X., and Meng, D · 2023
Later among the works it cites.
Mistral 7b, 2023
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Later among the works it cites.
Parameter-efficient orthogonal finetuning via butterfly factorization, 2023
Liu, W., Qiu, Z., Feng, Y., Xiu, Y., Xue, Y., Yu, L., Feng, H., Liu, Z., Heo, J., Peng, S., Wen, Y., Black, M. J., Weller, A., and Schölkopf, B · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nan, L., Radev, D., Zhang, R., Rau, A., Sivaprasad, A., Hsieh, C., Tang, X., Vyas, A., Verma, N., Krishna, P., Liu, Y., Irwanto, N., Pan, J., Rahman, F., Zaidi, A., Mutuma, M., Tarabar, Y., Gupta, A., Yu, T., Tan, Y. C., Lin, X. V., Xiong, C., Socher, R., and Rajani, N. F · 2021
Cited alongside, same era.
Low-rank matrix recovery with scaled subgradient methods: Fast and robust convergence without the condition number
Tong, T., Ma, C., and Chi, Y · 2021
Cited alongside, same era.
An image is worth one word: Personalizing text-to-image generation using textual inversion, 2022
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-Or, D · 2022
Cited alongside, same era.
Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps, 2022
Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J · 2022
Cited alongside, same era.
Fast convex optimization for two-layer relu networks: Equivalent model classes and cone decompositions, 2022
Mishkin, A., Sahiner, A., and Pilanci, M · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Scaling and scalability: Provable nonconvex low-rank tensor estimation from incomplete measurements, 2022
Tong, T., Ma, C., Prater-Bennette, A., Tripp, E., and Chi, Y · 2022
Cited alongside, same era.
Provably accelerating ill-conditioned low-rank estimation via scaled gradient descent, even with overparameterization, 2023
Ma, C., Xu, X., Tong, T., and Chi, Y · 2023
Later among the works it cites.
Controlling text-to-image diffusion by orthogonal finetuning, 2023
Qiu, Z., Liu, W., Feng, H., Xue, Y., Feng, Y., Liu, Z., Zhang, D., Weller, A., and Schölkopf, B · 2023
Later among the works it cites.
Low-rank adaptation for fast text-to-image diffusion fine-tuning
Ryu, S · 2023
Later among the works it cites.
Dylora: Parameter efficient tuning of pre-trained models using dynamic search-free low-rank adaptation, 2023
Valipour, M., Rezagholizadeh, M., Kobyzev, I., and Ghodsi, A · 2023
Later among the works it cites.
Delta-lora: Fine-tuning high-rank parameters with the delta of low-rank matrices, 2023
Zi, B., Qi, X., Wang, L., Wang, J., Wong, K.-F., and Zhang, L · 2023
Later among the works it cites.
Lora+: Efficient low rank adaptation of large models, 2024
Hayou, S., Ghosh, N., and Yu, B · 2024
Closest in time.
Mistral fine-tuning example, 2024
Labonne, M · 2024
Closest in time.
Mistral-7b announcement, 2023
Mistral AI team · 2024
Closest in time.