Fetching the paper…
Reading the bibliography…
Low-rank adaptation (LoRA) has become the standard approach for parameter-efficient fine-tuning of large language models (LLM), but our theoretical understanding of LoRA has been limited.
Extensions of lipschitz maps into a hilbert space
Johnson, W. and Lindenstrauss, J · 1984
Earlier work this paper cites.
Introduction to Optimization
Polyak, B. T · 1987
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G · 1989
Earlier work this paper cites.
On the method of bounded differences
McDiarmid, C. et al · 1989
Earlier work this paper cites.
Universal approximation of an unknown mapping and its derivatives using multilayer feedforward networks
Hornik, K., Stinchcombe, M., and White, H · 1990
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Barron, A. R · 1993
Earlier work this paper cites.
Problems of distance geometry and convex properties of quadratic maps
Barvinok, A. I · 1995
Earlier work this paper cites.
Critical points of matrix least squares distance functions
Helmke, U. and Shayman, M. A · 1995
Earlier work this paper cites.
On the rank of extreme matrices in semidefinite programs and the multiplicity of optimal eigenvalues
Pataki, G · 1998
Earlier work this paper cites.
Rademacher processes and bounding the risk of function learning
Koltchinskii, V. and Panchenko, D · 2000
Earlier work this paper cites.
The geometry of semidefinite programming
Pataki, G · 2000
Earlier work this paper cites.
Rademacher and gaussian complexities: risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2002
Earlier work this paper cites.
Model selection and error estimation
Bartlett, P. L., Boucheron, S., and Lugosi, G · 2002
Earlier work this paper cites.
Stability and generalization
Bousquet, O. and Elisseeff, A · 2002
Earlier work this paper cites.
A nonlinear programming algorithm for solving semidefinite programs via low-rank factorization
Burer, S. and Monteiro, R. D · 2003
Earlier work this paper cites.
Local rademacher complexities
Bartlett, P. L., Bousquet, O., and Mendelson, S · 2005
Earlier work this paper cites.
Convex sparse matrix factorizations
Bach, F., Mairal, J., and Ponce, J · 2008
Earlier work this paper cites.
Fast rates for regularized objectives
Sridharan, K., Shalev-Shwartz, S., and Srebro, N · 2008
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Recht, B., Fazel, M., and Parrilo, P. A · 2010
Earlier work this paper cites.
On the expressive power of deep architectures
Bengio, Y. and Delalleau, O · 2011
Earlier work this paper cites.
Shallow vs. deep sum-product networks
Delalleau, O. and Bengio, Y · 2011
Earlier work this paper cites.
Unifying nuclear norm and bilinear factorization approaches for low-rank matrix decomposition
Cabral, R., De la Torre, F., Costeira, J. P., and Bernardino, A · 2013
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Structured low-rank matrix factorization: optimality, algorithm, and applications to image processing
Haeffele, B., Young, E., and Vidal, R · 2014
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Earlier work this paper cites.
The non-convex Burer–Monteiro approach works on smooth semidefinite programs
Boumal, N., Voroninski, V., and Bandeira, A · 2016
Cited alongside, same era.
Matrix completion has no spurious local minimum
Ge, R., Lee, J. D., and Ma, T · 2016
Cited alongside, same era.
Train faster, generalize better: stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B · 2016
Cited alongside, same era.
A vector-contraction inequality for rademacher complexities
Maurer, A · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M. J · 2017
Cited alongside, same era.
Neural networks are convex regularizers: exact polynomial-time convex optimization formulations for two-layer networks
Pilanci, M. and Ergen, T · 2020
Later among the works it cites.
Why are adaptive methods good for attention models?
Zhang, J., Karimireddy, S. P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S · 2020
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2020
Later among the works it cites.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Aghajanyan, A., Zettlemoyer, L., and Gupta, S · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Cited alongside, same era.
Gradient descent can take exponential time to escape saddle points
Du, S. S., Jin, C., Lee, J. D., Jordan, M. I., Singh, A., and Poczos, B · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I · 2017
Cited alongside, same era.
The expressive power of neural networks: a view from the width
Lu, Z., Pu, H., Wang, F., Hu, Z., and Wang, L · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Towards understanding generalization of deep learning: perspective of loss landscapes
Wu, L., Zhu, Z., et al · 2017
Cited alongside, same era.
Making pre-trained language models better few-shot learners
Gao, T., Fisch, A., and Chen, D · 2021
Later among the works it cites.
LoRA: low-rank adaptation of large language models
Hu, E. J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al · 2021
Later among the works it cites.
Exploiting cloze questions for few shot text classification and natural language inference
Schick, T. and Schütze, H · 2021
Later among the works it cites.
Superb: Speech processing universal performance benchmark
Yang, S.-w., Chi, P.-H., Chuang, Y.-S., Lai, C.-I. J., Lakhotia, K., Lin, Y. Y., Liu, A. T., Shi, J., Chang, X., Lin, G.-T., et al · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2021
Later among the works it cites.
Learning Theory from First Principles
Bach, F · 2023
Later among the works it cites.
Transformers learn through gradual rank increase
Boix-Adsera, E., Littwin, E., Abbe, E., Bengio, S., and Susskind, J · 2023
Later among the works it cites.
LoRA can replace time and class embeddings in diffusion probabilistic models
Choi, J. Y., Park, J., Park, I., Cho, J., No, A., and Ryu, E. K · 2023
Later among the works it cites.
QLoRA: efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Later among the works it cites.
Minimum width of leaky-relu neural networks for uniform universal approximation
Duan, Y., Ji, G., Cai, Y., et al · 2023
Later among the works it cites.
On the effectiveness of parameter-efficient fine-tuning
Fu, Z., Yang, H., So, A. M.-C., Lam, W., Bing, L., and Collier, N · 2023
Later among the works it cites.
Looped transformers as programmable computers
Giannou, A., Rajput, S., Sohn, J.-y., Lee, K., Lee, J. D., and Papailiopoulos, D · 2023
Later among the works it cites.
Understanding incremental learning of gradient descent: A fine-grained analysis of matrix sensing
Jin, J., Li, Z., Lyu, K., Du, S. S., and Lee, J. D · 2023
Later among the works it cites.
ReLoRA: high-rank training through low-rank updates
Lialin, V., Muckatira, S., Shivagunde, N., and Rumshisky, A · 2023
Later among the works it cites.
A kernel-based view of language model fine-tuning
Malladi, S., Wettig, A., Yu, D., Chen, D., and Arora, S · 2023
Later among the works it cites.
Low-rank adaptation for fast text-to-image diffusion fine-tuning, 2023
Ryu, S · 2023
Later among the works it cites.
Continual diffusion: continual customization of text-to-image diffusion with c-lora
Smith, J. S., Hsu, Y.-C., Zhang, L., Hua, T., Kira, Z., Shen, Y., and Jin, H · 2023
Later among the works it cites.
Navigating text-to-image customization: from LyCORIS fine-tuning to model evaluation
Yeh, S.-Y., Hsieh, Y.-G., Gao, Z., Yang, B. B., Oh, G., and Gong, Y · 2024
Closest in time.
The expressive power of low-rank adaptation
Zeng, Y. and Lee, K · 2024
Closest in time.