Fetching the paper…
Reading the bibliography…
In the evolving landscape of natural language processing (NLP), fine-tuning pre-trained Large Language Models (LLMs) with first-order (FO) optimizers like SGD and Adam has become standard.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Backpropagation and stochastic gradient descent method
Amari, S.-i · 1993
Earlier work this paper cites.
Online convex optimization in the bandit setting: gradient descent without a gradient
Flaxman, A. D., Kalai, A. T., and McMahan, H. B · 2004
Earlier work this paper cites.
Online convex optimization in the bandit setting: Gradient descent without a gradient
Flaxman, A. D., Kalai, A. T., and McMahan, H. B · 2005
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Roemmele, M., Bejan, C. A., and Gordon, A. S · 2011
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Optimal rates for zero-order convex optimization: The power of two function evaluations
Duchi, J. C., Jordan, M. I., Wainwright, M. J., and Wibisono, A · 2015
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models
Chen, P.-Y., Zhang, H., Sharma, Y., Yi, J., and Hsieh, C.-J · 2017
Earlier work this paper cites.
Decoupled neural interfaces using synthetic gradients
Jaderberg, M., Czarnecki, W. M., Osindero, S., Vinyals, O., Graves, A., Silver, D., and Kavukcuoglu, K · 2017
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Nesterov, Y. and Spokoiny, V · 2017
Earlier work this paper cites.
Stochastic zeroth-order optimization in high dimensions
Wang, Y., Du, S., Balakrishnan, S., and Singh, A · 2017
Earlier work this paper cites.
Zeroth-order (non)-convex stochastic optimization via conditional gradient and gradient updates
Balasubramanian, K. and Ghadimi, S · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Universal decision-based black-box perturbations: Breaking security-through-obscurity defenses
Hogan, T. A. and Kailkhura, B · 2018
Earlier work this paper cites.
Black-box adversarial attacks with limited queries and information
Ilyas, A., Engstrom, L., Athalye, A., and Lin, J · 2018
Earlier work this paper cites.
Looking beyond the surface:a challenge set for reading comprehension over multiple sentences
Khashabi, D., Chaturvedi, S., Roth, M., Upadhyay, S., and Roth, D · 2018
Earlier work this paper cites.
Zeroth-order stochastic variance reduction for nonconvex optimization
Liu, S., Kailkhura, B., Chen, P.-Y., Ting, P., Chang, S., and Amini, L · 2018
Earlier work this paper cites.
Hessian-aware zeroth-order optimization for black-box adversarial attack
Ye, H., Huang, Z., Fang, C., Li, C. J., and Zhang, T · 2018
Earlier work this paper cites.
Zo-adamm: Zeroth-order adaptive momentum method for black-box optimization
Chen, X., Liu, S., Xu, K., Li, X., Lin, X., Hong, M., and Cox, D · 2019
Earlier work this paper cites.
Model agnostic contrastive explanations for structured data
Dhurandhar, A., Pedapati, T., Balakrishnan, A., Chen, P.-Y., Shanmugam, K., and Puri, R · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
Improving gradient estimation in evolutionary strategies with past descent directions
Meier, F., Mujika, A., Gauy, M. M., and Steger, A · 2019
Earlier work this paper cites.
Training neural networks with local error signals
Nøkland, A. and Eidnes, L. H · 2019
Earlier work this paper cites.
Autozoom: Autoencoder-based zeroth order optimization method for attacking black-box neural networks
Tu, C.-C., Ting, P., Chen, P.-Y., Liu, S., Zhang, H., Yi, J., Hsieh, C.-J., and Cheng, S.-M · 2019
Cited alongside, same era.
Contrasting exploration in parameter and action space: A zeroth-order optimization perspective
Vemula, A., Sun, W., and Bagnell, J · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Cited alongside, same era.
On the design of black-box adversarial examples by leveraging gradient-free optimization and operator splitting method
Zhao, P., Liu, S., Chen, P.-Y., Hoang, N., Xu, K., Kailkhura, B., and Lin, X · 2019
Cited alongside, same era.
The lottery ticket hypothesis for pre-trained bert networks
Chen, T., Frankle, J., Chang, S., Liu, S., Zhang, Y., Wang, Z., and Carbin, M · 2020
Cited alongside, same era.
Clip-tuning: Towards derivative-free prompt learning with a mixture of rewards, 2022
Chai, Y., Wang, S., Sun, Y., Tian, H., Wu, H., and Wang, H · 2022
Later among the works it cites.
Adaptformer: Adapting vision transformers for scalable visual recognition
Chen, S., Ge, C., Tong, Z., Wang, J., Song, Y., Wang, J., and Luo, P · 2022
Later among the works it cites.
Rlprompt: Optimizing discrete text prompts with reinforcement learning, 2022
Deng, M., Wang, J., Hsieh, C.-P., Wang, Y., Guo, H., Shu, T., Song, M., Xing, E. P., and Hu, Z · 2022
Later among the works it cites.
The forward-forward algorithm: Some preliminary investigations
Hinton, G · 2022
Later among the works it cites.
Optimizing molecules using efficient queries from property evaluations
Hoffman, S. C., Chenthamarakshan, V., Wadhawan, K., Chen, P.-Y., and Das, P · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gao, T., Fisch, A., and Chen, D · 2020
Cited alongside, same era.
Exploring versatile generative language model via parameter-efficient transfer learning, 2020
Lin, Z., Madotto, A., and Fung, P · 2020
Cited alongside, same era.
A primer on zeroth-order optimization in signal processing and machine learning: Principals, recent advances, and applications
Liu, S., Chen, P.-Y., Kailkhura, B., Zhang, G., Hero III, A. O., and Varshney, P. K · 2020
Cited alongside, same era.
Sparse perturbations for improved convergence in stochastic zeroth-order optimization
Ohta, M., Berger, N., Sokolov, A., and Riezler, S · 2020
Cited alongside, same era.
Adapterhub: A framework for adapting transformers
Pfeiffer, J., Rücklé, A., Poth, C., Kamath, A., Vulić, I., Ruder, S., Cho, K., and Gurevych, I · 2020
Cited alongside, same era.
Transfer learning without knowing: Reprogramming black-box machine learning models with scarce data and limited resources
Tsai, Y.-Y., Chen, P.-Y., and Ho, T.-Y · 2020
Cited alongside, same era.
A zeroth-order block coordinate descent algorithm for huge-scale black-box optimization
Cai, H., Lou, Y., McKenzie, D., and Yin, W · 2021
Cited alongside, same era.
Accelerated zeroth-order and first-order momentum methods from mini to minimax optimization
Huang, F., Gao, S., Pei, J., and Huang, H · 2022
Later among the works it cites.
P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks
Liu, X., Ji, K., Fu, Y., Tam, W., Du, Z., Yang, Z., and Tang, J · 2022
Later among the works it cites.
Accelerating dnn training with structured data gradient pruning
McDanel, B., Dinh, H., and Magallanes, J · 2022
Later among the works it cites.
Scaling forward gradient with local losses
Ren, M., Kornblith, S., Liao, R., and Hinton, G · 2022
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization, 2022
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., Dey, M., Bari, M. S., Xu, C., Thakker, U., Sharma, S. S., Szczechla, E., Kim, T., Chhablani, G., Nayak, N., Datta, D., Chang, J., Jiang, M. T.-J., Wang, H., Manica, M., Shen, S., Yong, Z. X., Pandey, H., Bawden, R., Wang, T., Neeraj, T., Rozen, J., Sharma, A., Santilli, A., Fevry, T., Fries, J. A., Teehan, R., Bers, T., Biderman, S., Gao, L., Wolf, T., and Rush, A. M · 2022
Later among the works it cites.
Zeroth-order optimization with trajectory-informed derivative estimation
Shu, Y., Dai, Z., Sng, W., Verma, A., Jaillet, P., and Low, B. K. H · 2022
Later among the works it cites.
Zeroth-order sciml: Non-intrusive integration of scientific software with deep learning
Tsaknakis, I., Kailkhura, B., Liu, S., Loveland, D., Diffenderfer, J., Hiszpanski, A. M., and Hong, M · 2022
Later among the works it cites.
Zarts: On zero-order optimization for neural architecture search
Wang, X., Guo, W., Su, J., Yang, X., and Yan, J · 2022
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Later among the works it cites.
BPipe: Memory-balanced pipeline parallelism for training large language models
Kim, T., Kim, H., Yu, G.-I., and Chun, B.-G · 2023
Later among the works it cites.
Towards efficient visual adaption via structural re-parameterization
Luo, G., Huang, M., Zhou, Y., Sun, X., Jiang, G., Wang, Z., and Ji, R · 2023
Later among the works it cites.
Fine-tuning language models with just forward passes
Malladi, S., Gao, T., Nichani, E., Damian, A., Lee, J. D., Chen, D., and Arora, S · 2023
Later among the works it cites.
Grips: Gradient-free, edit-based instruction search for prompting large language models, 2023
Prasad, A., Hase, P., Zhou, X., and Bansal, M · 2023
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2023
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2023
Later among the works it cites.
Singhal, U., Cheung, B., Chandra, K., Ragan-Kelley, J., Tenenbaum, J. B., Poggio, T. A., and Yu, S. X · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Certified zeroth-order black-box defense with robust unet denoiser
Verma, A., Bangar, S., Subramanyam, A., Lal, N., Shah, R. R., and Satoh, S · 2023
Later among the works it cites.
Low-variance gradient estimation in unrolled computation graphs with es-single
Vicol, P., Kolter, Z., and Swersky, K · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.
On-device training: A first overview on existing systems, 2023
Zhu, S., Voigt, T., Ko, J., and Rahimian, F · 2023
Later among the works it cites.
Deepzero: Scaling up zeroth-order optimization for deep model training
Chen, A., Zhang, Y., Jia, J., Diffenderfer, J., Liu, J., Parasyris, K., Zhang, Y., Zhang, Z., Kailkhura, B., and Liu, S · 2024
Closest in time.