Fetching the paper…
Reading the bibliography…
Zeroth-order optimization (ZO) has demonstrated remarkable promise in efficient fine-tuning tasks for Large Language Models (LLMs).
Roberta: A robustly optimized bert pretraining approach. arxiv [preprint](2019)
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
The expression of a tensor or a polyadic as a sum of products
Hitchcock, F. L · 1927
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
Spall, J. C · 1992
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models
Chen, P.-Y., Zhang, H., Sharma, Y., Yi, J., and Hsieh, C.-J · 2017
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Nesterov, Y. and Spokoiny, V · 2017
Earlier work this paper cites.
Stochastic zeroth-order optimization in high dimensions
Wang, Y., Du, S., Balakrishnan, S., and Singh, A · 2018
Earlier work this paper cites.
Gradientless descent: High-dimensional zeroth-order optimization
Golovin, D., Karro, J., Kochanski, G., Lee, C., Song, X., and Zhang, Q · 2019
Earlier work this paper cites.
Zone: Zeroth-order nonconvex multiagent optimization over networks
Hajinezhad, D., Hong, M., and Garcia, A · 2019
Earlier work this paper cites.
Improved zeroth-order variance reduced algorithms and analysis for nonconvex optimization
Ji, K., Wang, Z., Zhou, Y., and Liang, Y · 2019
Earlier work this paper cites.
Autozoom: Autoencoder-based zeroth order optimization method for attacking black-box neural networks
Tu, C.-C., Ting, P., Chen, P.-Y., Liu, S., Zhang, H., Yi, J., Hsieh, C.-J., and Cheng, S.-M · 2019
Earlier work this paper cites.
Contrasting exploration in parameter and action space: A zeroth-order optimization perspective
Vemula, A., Sun, W., and Bagnell, J · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
A primer on zeroth-order optimization in signal processing and machine learning: Principals, recent advances, and applications
Liu, S., Chen, P.-Y., Kailkhura, B., Zhang, G., Hero III, A. O., and Varshney, P. K · 2020
Earlier work this paper cites.
Distributed zero-order algorithms for nonconvex multiagent optimization
Tang, Y., Zhang, J., and Li, N · 2020
Earlier work this paper cites.
Towards query-efficient black-box adversary with zeroth-order natural gradient descent
Zhao, P., Chen, P.-Y., Wang, S., and Lin, X · 2020
Earlier work this paper cites.
On the convergence of prior-guided zeroth-order optimization algorithms
Cheng, S., Wu, G., and Zhu, J · 2021
Cited alongside, same era.
Privacy-preserved distributed learning with zeroth-order optimization
Gratton, C., Venkategowda, N. K., Arablouei, R., and Werner, S · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Maze: Data-free model stealing attack using zeroth-order gradient estimation
Kariyappa, S., Prakash, A., and Qureshi, M. K · 2021
Cited alongside, same era.
Linear convergence of first-and zeroth-order primal–dual algorithms for distributed nonconvex optimization
Yi, X., Zhang, S., Yang, T., Chai, T., and Johansson, K. H · 2021
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Cuttlefish: Low-rank model training without all the tuning
Wang, H., Agarwal, S., Tanaka, Y., Xing, E., Papailiopoulos, D., et al · 2023
Later among the works it cites.
Dpzero: dimension-independent and differentially private zeroth-order optimization
Zhang, L., Thekumparampil, K. K., Oh, S., and He, N · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2023
Later among the works it cites.
Enhancing zeroth-order fine-tuning for language models with low-rank structures
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Communication-efficient stochastic zeroth-order optimization for federated learning
Fang, W., Yu, Z., Jiang, Y., Shi, Y., Jones, C. N., and Zhou, Y · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Cited alongside, same era.
Black-box generalization: Stability of zeroth-order learning
Nikolakakis, K., Haddadpour, F., Kalogerias, D., and Karbasi, A · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Deepzero: Scaling up zeroth-order optimization for deep model training
Chen, A., Zhang, Y., Jia, J., Diffenderfer, J., Liu, J., Parasyris, K., Zhang, Y., Zhang, Z., Kailkhura, B., and Liu, S · 2023
Cited alongside, same era.
Fine-tuning language models with just forward passes
Malladi, S., Gao, T., Nichani, E., Damian, A., Lee, J. D., Chen, D., and Arora, S · 2023
Cited alongside, same era.
Chen, Y., Zhang, Y., Cao, L., Yuan, K., and Wen, Z · 2024
Later among the works it cites.
Variance-reduced zeroth-order methods for fine-tuning language models
Gautam, T., Park, Y., Zhou, H., Raman, P., and Ha, W · 2024
Later among the works it cites.
Zeroth-order fine-tuning of llms with extreme sparsity
Guo, W., Long, J., Zeng, Y., Liu, Z., Yang, X., Ran, Y., Gardner, J. R., Bastani, O., De Sa, C., Yu, X., et al · 2024
Later among the works it cites.
On the inherent privacy of two point zeroth order projected gradient descent
Gupta, D., Razaviyayn, M., and Sharan, V · 2024
Later among the works it cites.
Nonconvex zeroth-order stochastic admm methods with lower function query complexity
Huang, F., Gao, S., Pei, J., and Huang, H · 2024
Later among the works it cites.
From galore to welore: How low-rank weights non-uniformly emerge from low-rank gradients
Jaiswal, A., Yin, L., Zhang, Z., Liu, S., Zhao, J., Tian, Y., and Wang, Z · 2024
Later among the works it cites.
Zo-adamu optimizer: Adapting perturbation by the momentum and uncertainty in zeroth-order optimization
Jiang, S., Chen, Q., Pan, Y., Xiang, Y., Lin, Y., Wu, X., Liu, C., and Song, X · 2024
Later among the works it cites.
Efficient zeroth-order proximal stochastic method for nonconvex nonsmooth black-box problems
Kazemi, E. and Wang, L · 2024
Later among the works it cites.
Sparse mezo: Less parameters for better performance in zeroth-order llm fine-tuning
Liu, Y., Zhu, Z., Gong, C., Cheng, M., Hsieh, C.-J., and You, Y · 2024
Later among the works it cites.
An optimal structured zeroth-order algorithm for non-smooth optimization
Rando, M., Molinari, C., Rosasco, L., and Villa, S · 2024
Later among the works it cites.
Yang, Y., Zhen, K., Banijamal, E., Mouchtaris, A., and Zhang, Z · 2024
Later among the works it cites.
Subzero: Random subspace zeroth-order optimization for memory-efficient llm fine-tuning
Yu, Z., Zhou, P., Wang, S., Li, J., and Huang, H · 2024
Later among the works it cites.