Fetching the paper…
Reading the bibliography…
Fine-tuning large language models (LLMs) with classic first-order optimizers entails prohibitive GPU memory due to the backpropagation process.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
The Convergence of a Class of Double-rank Minimization Algorithms 1. General Considerations
C. G. BROYDEN · 1970
Earlier work this paper cites.
The moments of products of quadratic forms in normal variables
Jan R Magnus et al · 1978
Earlier work this paper cites.
Inexact newton methods
Ron S. Dembo, Stanley C. Eisenstat, and Trond Steihaug · 1982
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
J.C. Spall · 1992
Earlier work this paper cites.
A one-measurement form of simultaneous perturbation stochastic approximation
James C. Spall · 1997
Earlier work this paper cites.
Improving the convergence of the backpropagation algorithm using learning rate adaptation methods
G. D. Magoulas, M. N. Vrahatis, and G. S. Androulakis · 1999
Earlier work this paper cites.
Trust-Region Methods
Andrew R. Conn, Nicholas I. M. Gould, and Philippe L. Toint · 2000
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Yurii Nesterov and B. T. Polyak · 2006
Earlier work this paper cites.
Information-theoretic lower bounds on the oracle complexity of convex optimization
Alekh Agarwal, Martin J Wainwright, Peter Bartlett, and Pradeep Ravikumar · 2009
Earlier work this paper cites.
Algorithm for stochastic approximation with trial input perturbation in the nonstationary problem of optimization
A. T. Vakhitov, O. N. Granichin, and L. S. Gurevich · 2009
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Information-based complexity, feedback and dynamics in convex programming
Maxim Raginsky and Alexander Rakhlin · 2011
Earlier work this paper cites.
Galore: Memory-efficient llm training by gradient low-rank projection, 2024b
Jiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang, Anima Anandkumar, and Yuandong Tian · 2011
Earlier work this paper cites.
Query complexity of derivative-free optimization
Kevin G Jamieson, Robert Nowak, and Ben Recht · 2012
Earlier work this paper cites.
Adadelta: An adaptive learning rate method, 2012
Matthew D. Zeiler · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Revisiting natural gradient for deep networks, 2014
Razvan Pascanu and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models
Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh · 2017
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond, 2017
Levent Sagun, Leon Bottou, and Yann LeCun · 2017
Cited alongside, same era.
Fast approximate natural gradient descent in a kronecker-factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Cited alongside, same era.
Gradient-free multi-agent nonconvex nonsmooth optimization
Davood Hajinezhad and Michael M. Zavlanos · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
Inexact non-convex newton-type methods, 2018
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Later among the works it cites.
Adahessian: An adaptive second order optimizer for machine learning
Zhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa, Kurt Keutzer, and Michael Mahoney · 2021
Later among the works it cites.
A new one-point residual-feedback oracle for black-box learning and control
Yan Zhang, Yi Zhou, Kaiyi Ji, and Michael M. Zavlanos · 2021
Later among the works it cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms, 2023
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Later among the works it cites.
Textbooks are all you need, 2023
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhewei Yao, Peng Xu, Farbod Roosta-Khorasani, and Michael W. Mahoney · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Newton-type methods for non-convex optimization under inexact hessian information, 2019
Peng Xu, Fred Roosta, and Michael W. Mahoney · 2019
Cited alongside, same era.
Hessian-aware zeroth-order optimization for black-box adversarial attack, 2019
Haishan Ye, Zhichao Huang, Cong Fang, Chris Junchi Li, and Tong Zhang · 2019
Cited alongside, same era.
Later among the works it cites.
Zo-adamu optimizer: Adapting perturbation by the momentum and uncertainty in zeroth-order optimization, 2023
Shuoran Jiang, Qingcai Chen, Youchen Pan, Yang Xiang, Yukang Lin, Xiangping Wu, Chuanyi Liu, and Xiaobao Song · 2023
Later among the works it cites.
Textbooks are all you need ii: phi-1.5 technical report, 2023
Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, and Yin Tat Lee · 2023
Later among the works it cites.
Sophia: A scalable stochastic second-order optimizer for language model pre-training
Hong Liu and Zhiyuan Li · 2023
Later among the works it cites.
Fine-tuning language models with just forward passes
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D. Lee, Danqi Chen, and Sanjeev Arora · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Mirror natural evolution strategies, 2023
Haishan Ye · 2023
Later among the works it cites.
Eva: Practical second-order optimization with kronecker-vectorized approximation
Lin Zhang, Shaohuai Shi, and Bo Li · 2023
Later among the works it cites.
Enhancing zeroth-order fine-tuning for language models with low-rank structures
Yiming Chen, Yuan Zhang, Liyuan Cao, Kun Yuan, and Zaiwen Wen · 2024
Closest in time.
Zeroth-order fine-tuning of llms with extreme sparsity
Wentao Guo, Jikai Long, Yimeng Zeng, Zirui Liu, Xinyu Yang, Yide Ran, Jacob R Gardner, Osbert Bastani, Christopher De Sa, Xiaodong Yu, et al · 2024
Closest in time.
Sparse mezo: Less parameters for better performance in zeroth-order llm fine-tuning
Yong Liu, Zirui Zhu, Chaoyu Gong, Minhao Cheng, Cho-Jui Hsieh, and Yang You · 2024
Closest in time.
Lisa: Layerwise importance sampling for memory-efficient large language model fine-tuning, 2024
Rui Pan, Xiang Liu, Shizhe Diao, Renjie Pi, Jipeng Zhang, Chi Han, and Tong Zhang · 2024
Closest in time.
Effectively learning from data and generating data in differentially private machine learning
Xinyu Tang · 2024
Closest in time.
Revisiting zeroth-order optimization for memory-efficient LLM fine-tuning: A benchmark
Yihua Zhang, Pingzhi Li, Junyuan Hong, Jiaxiang Li, Yimeng Zhang, Wenqing Zheng, Pin-Yu Chen, Jason D. Lee, Wotao Yin, Mingyi Hong, Zhangyang Wang, Sijia Liu, and Tianlong Chen · 2024
Closest in time.
Yan Sun, Tiansheng Huang, Liang Ding, Li Shen, and Dacheng Tao · 2025
Closest in time.
Harmony in divergence: Towards fast, accurate, and memory-efficient zeroth-order llm fine-tuning
Qitao Tan, Jun Liu, Zheng Zhan, Caiwei Ding, Yanzhi Wang, Jin Lu, and Geng Yuan · 2025
Closest in time.