Fetching the paper…
Reading the bibliography…
In this paper, we study the role of initialization in Low Rank Adaptation (LoRA) as originally introduced in Hu et al.
“RoBERTa: A Robustly Optimized BERT Pretraining Approach”, 2019
Yinhan Liu et al · 1907
Earlier work this paper cites.
“Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning”
Haokun Liu et al · 1965
Earlier work this paper cites.
“Efficient backprop”
Yann LeCun, Léon Bottou, Genevieve Orr and Klaus-Robert Müller · 2002
Earlier work this paper cites.
“A theory of transfer learning with applications to active learning”
Liu Yang, Steve Hanneke and Jaime Carbonell · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization”
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
“Deep residual learning for image recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“Deep Information Propagation”
S.S. Schoenholz, J. Gilmer, S. Ganguli and J. Sohl-Dickstein · 2017
Earlier work this paper cites.
“Deep Information Propagation”, 2017
Samuel. Schoenholz, Justin Gilmer, Surya Ganguli and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
“Improving language understanding by generative pre-training”
Alec Radford, Karthik Narasimhan, Tim Salimans and Ilya Sutskever · 2018
Earlier work this paper cites.
“GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding”, 2018
Alex Wang et al · 2018
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“On the Impact of the Activation function on Deep Neural Networks Training”
Soufiane Hayou, Arnaud Doucet and Judith Rousseau · 2019
Earlier work this paper cites.
“Parameter-efficient transfer learning for NLP”
Neil Houlsby et al · 2019
Earlier work this paper cites.
G. Yang · 2019
Earlier work this paper cites.
“Scaling laws for neural language models”
Jared Kaplan et al · 2020
Earlier work this paper cites.
“Tensor programs iii: Neural matrix laws”
Greg Yang · 2020
Earlier work this paper cites.
“Training verifiers to solve math word problems”
Karl Cobbe et al · 2021
Earlier work this paper cites.
“Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability”
Jeremy Cohen et al · 2021
Cited alongside, same era.
“Stable ResNet”
Soufiane Hayou et al · 2021
Cited alongside, same era.
“LoRA: Low-Rank Adaptation of Large Language Models”
Edward. Hu et al · 2021
Cited alongside, same era.
“The power of scale for parameter-efficient prompt tuning”
Brian Lester, Rami Al-Rfou and Noah Constant · 2021
Cited alongside, same era.
“Tensor programs iv: Feature learning in infinite-width neural networks”
Greg Yang and Edward Hu · 2021
Cited alongside, same era.
“Improved baselines with visual instruction tuning”
Haotian Liu, Chunyuan Li, Yuheng Li and Yong Lee · 2023
Later among the works it cites.
“The flan collection: Designing data and methods for effective instruction tuning”
Shayne Longpre et al · 2023
Later among the works it cites.
“The Shaped Transformer: Attention Models in the Infinite Depth-and-Width Limit”, 2023
Lorenzo Noci et al · 2023
Later among the works it cites.
“Llama 2: Open Foundation and Fine-Tuned Chat Models”
Hugo Touvron et al · 2023
Later among the works it cites.
“How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jordan Hoffmann et al · 2022
Cited alongside, same era.
“The Neural Covariance SDE: Shaped Infinite Depth-and-Width Networks at Initialization”
Mufan Li, Mihai Nica and Dan Roy · 2022
Cited alongside, same era.
“Emergent abilities of large language models”
Jason Wei et al · 2022
Cited alongside, same era.
“Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer”
Greg Yang et al · 2022
Cited alongside, same era.
“QLoRA: Efficient Finetuning of Quantized LLMs”
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman and Luke Zettlemoyer · 2023
Cited alongside, same era.
“On the infinite-depth limit of finite-width neural networks”
Soufiane Hayou · 2023
Cited alongside, same era.
“Width and Depth Limits Commute in Residual Networks”
Soufiane Hayou and Greg Yang · 2023
Cited alongside, same era.
Yizhong Wang et al · 2023
Later among the works it cites.
“Tensor programs ivb: Adaptive optimization in the infinite-width limit”
Greg Yang and Etai Littwin · 2023
Later among the works it cites.
“Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks”
Greg Yang, Dingli Yu, Chen Zhu and Soufiane Hayou · 2023
Later among the works it cites.
“Lora-fa: Memory-efficient low-rank adaptation for large language models fine-tuning”
Longteng Zhang et al · 2023
Later among the works it cites.
“LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters”
Klaudia Baazy, Mohammadreza Banaei, Karl Aberer and Jacek Tabor · 2024
Closest in time.
“LoRA+: Efficient Low Rank Adaptation of Large Models”, 2024
Soufiane Hayou, Nikhil Ghosh and Bin Yu · 2024
Closest in time.
“MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning”
Ting Jiang et al · 2024
Closest in time.
“VB-LoRA: Extreme Parameter Efficient Fine-Tuning with Vector Banks”
Yang Li, Shaobo Han and Shihao Ji · 2024
Closest in time.
“DoRA: Weight-Decomposed Low-Rank Adaptation”
Shih-Yang Liu et al · 2024
Closest in time.
“PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models”
Fanxu Meng, Zhaohui Wang and Muhan Zhang · 2024
Closest in time.
“Tinyllama: An open-source small language model”
Peiyuan Zhang, Guangtao Zeng, Tianduo Wang and Wei Lu · 2024
Closest in time.
“Asymmetry in Low-Rank Adapters of Foundation Models”, 2024
Jiacheng Zhu et al · 2024
Closest in time.