Fetching the paper…
Reading the bibliography…
Zeroth-order (ZO) optimization has become a popular technique for solving machine learning (ML) problems when first-order (FO) information is difficult or impossible to obtain.
Numerical solution of a minimum problem
Enrico Fermi · 1952
Earlier work this paper cites.
Stochastic estimation of the maximum of a regression function
Jack Kiefer and Jacob Wolfowitz · 1952
Earlier work this paper cites.
A simplex method for function minimization
John A Nelder and Roger Mead · 1965
Earlier work this paper cites.
Learning characteristics of stochastic-gradient-descent algorithms: A general study, analysis, and critique
William A Gardner · 1984
Earlier work this paper cites.
On the convergence of the multidirectional search algorithm
Virginia Torczon · 1991
Earlier work this paper cites.
Weight perturbation: An optimal architecture and learning technique for analog vlsi feedforward and recurrent multilayer networks
Marwan Jabri and Barry Flower · 1992
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
James C Spall · 1992
Earlier work this paper cites.
Backpropagation and stochastic gradient descent method
Shun-ichi Amari · 1993
Earlier work this paper cites.
Genetic algorithms and machine learning
John J Grefenstette · 1993
Earlier work this paper cites.
Convex analysis and minimization algorithms
Jean-Baptiste Hiriart Urruty and Claude Lemaréchal · 1993
Earlier work this paper cites.
Backpropagation: The basic theory
David E Rumelhart, Richard Durbin, Richard Golden, and Yves Chauvin · 1995
Earlier work this paper cites.
The simplex gradient and noisy optimization problems
David Matthew Bortz and Carl Tim Kelley · 1998
Earlier work this paper cites.
Numerical optimization
Stephen Wright, Jorge Nocedal, et al · 1999
Earlier work this paper cites.
Trust region methods
Andrew R Conn, Nicholas IM Gould, and Philippe L Toint · 2000
Earlier work this paper cites.
Online convex optimization in the bandit setting: Gradient descent without a gradient
A. D. Flaxman, A. T. Kalai, and H. B. McMahan · 2005
Earlier work this paper cites.
Introduction to derivative-free optimization , volume 8
A. R. Conn, K. Scheinberg, and L. N. Vicente · 2009
Earlier work this paper cites.
Pswarm: a hybrid solver for linearly constrained global derivative-free optimization
A Ismael F Vaz and Luís Nunes Vicente · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Stochastic gradient descent tricks
Léon Bottou · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
Derivative-free optimization: a review of algorithms and comparison of software implementations
L. M. Rios and N. V. Sahinidis · 2013
Earlier work this paper cites.
On the complexity of bandit and derivative-free stochastic convex optimization
O. Shamir · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Random feedback weights support learning in deep neural networks
Timothy P Lillicrap, Daniel Cownden, Douglas B Tweed, and Colin J Akerman · 2014
Earlier work this paper cites.
Optimal rates for zero-order convex optimization: The power of two function evaluations
John C Duchi, Michael I Jordan, Martin J Wainwright, and Andre Wibisono · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Y. Nesterov and V. Spokoiny · 2015
Earlier work this paper cites.
Taking the human out of the loop: A review of bayesian optimization
Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P Adams, and Nando De Freitas · 2015
Earlier work this paper cites.
A comprehensive linear speedup analysis for asynchronous stochastic parallel optimization from zeroth-order to first-order
X. Lian, H. Zhang, C.-J. Hsieh, Y. Huang, and J. Liu · 2016
Earlier work this paper cites.
Direct feedback alignment provides learning in deep neural networks
Arild Nøkland · 2016
Earlier work this paper cites.
Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models
Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Decoupled neural interfaces using synthetic gradients
Max Jaderberg, Wojciech Marian Czarnecki, Simon Osindero, Oriol Vinyals, Alex Graves, David Silver, and Koray Kavukcuoglu · 2017
Cited alongside, same era.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Cited alongside, same era.
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny · 2017
Cited alongside, same era.
Stochastic zeroth-order optimization in high dimensions
Yining Wang, Simon Du, Sivaraman Balakrishnan, and Aarti Singh · 2017
Cited alongside, same era.
Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising
Kai Zhang, Wangmeng Zuo, Yunjin Chen, Deyu Meng, and Lei Zhang · 2017
Cited alongside, same era.
Transfer learning without knowing: Reprogramming black-box machine learning models with scarce data and limited resources
Yun-Yun Tsai, Pin-Yu Chen, and Tsung-Yi Ho · 2020
Later among the works it cites.
Solver-in-the-loop: Learning from differentiable physics to interact with iterative pde-solvers
Kiwon Um, Robert Brand, Yun Raymond Fei, Philipp Holl, and Nils Thuerey · 2020
Later among the works it cites.
Picking winning tickets before training by preserving gradient flow
Chaoqi Wang, Guodong Zhang, and Roger Grosse · 2020
Later among the works it cites.
A zeroth-order block coordinate descent algorithm for huge-scale black-box optimization
HanQin Cai, Yuchen Lou, Daniel McKenzie, and Wotao Yin · 2021
Later among the works it cites.
On the convergence of prior-guided zeroth-order optimization algorithms
Shuyu Cheng, Guoqiang Wu, and Jun Zhu · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zeroth-order (non)-convex stochastic optimization via conditional gradient and gradient updates
Krishnakumar Balasubramanian and Saeed Ghadimi · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Cited alongside, same era.
Universal decision-based black-box perturbations: Breaking security-through-obscurity defenses
Thomas A Hogan and Bhavya Kailkhura · 2018
Cited alongside, same era.
Error-gated hebbian rule: A local learning rule for principal and independent component analysis
Takuya Isomura and Taro Toyoizumi · 2018
Cited alongside, same era.
Snip: Single-shot network pruning based on connection sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip HS Torr · 2018
Cited alongside, same era.
Hessian-aware zeroth-order optimization for black-box adversarial attack
Haishan Ye, Zhichao Huang, Cong Fang, Chris Junchi Li, and Tong Zhang · 2018
Cited alongside, same era.
Imagenet training in minutes
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer · 2018
Cited alongside, same era.
Later among the works it cites.
James Diffenderfer and Bhavya Kailkhura · 2021
Later among the works it cites.
A winning hand: Compressing deep networks can improve out-of-distribution robustness
James Diffenderfer, Brian Bartoldson, Shreya Chaganti, Jize Zhang, and Bhavya Kailkhura · 2021
Later among the works it cites.
Sanity checks for lottery tickets: Does your winning ticket really win the jackpot?
Xiaolong Ma, Geng Yuan, Xuan Shen, Tianlong Chen, Xuxi Chen, Xiaohan Chen, Ning Liu, Minghai Qin, Sijia Liu, Zhangyang Wang, et al · 2021
Later among the works it cites.
Learning by directional gradient descent
David Silver, Anirudh Goyal, Ivo Danihelka, Matteo Hessel, and Hado van Hasselt · 2021
Later among the works it cites.
Gradients without backpropagation
Atılım Güneş Baydin, Barak A Pearlmutter, Don Syme, Frank Wood, and Philip Torr · 2022
Later among the works it cites.
Optimization without backpropagation
Gabriel Belouze · 2022
Later among the works it cites.
A theoretical and empirical comparison of gradient approximations in derivative-free optimization
Albert S Berahas, Liyuan Cao, Krzysztof Choromanski, and Katya Scheinberg · 2022
Later among the works it cites.
How to train your wide neural network without backprop: An input-weight alignment perspective
Akhilan Boopathy and Ila Fiete · 2022
Later among the works it cites.
Zeroth-order regularized optimization (zoro): Approximately sparse gradients and adaptive sampling
HanQin Cai, Daniel Mckenzie, Wotao Yin, and Zhenliang Zhang · 2022
Later among the works it cites.
Black-box prompt learning for pre-trained language models
Shizhe Diao, Zhichao Huang, Ruijia Xu, Xuechun Li, Yong Lin, Xiao Zhou, and Tong Zhang · 2022
Later among the works it cites.
Complex locomotion skill learning via differentiable physics
Yu Fang, Jiancheng Liu, Mingrui Zhang, Jiasheng Zhang, Yidong Ma, Minchen Li, Yuanming Hu, Chenfanfu Jiang, and Tiantian Liu · 2022
Later among the works it cites.
The forward-forward algorithm: Some preliminary investigations
Geoffrey Hinton · 2022
Later among the works it cites.
Softhebb: Bayesian inference in unsupervised hebbian soft winner-take-all networks
Timoleon Moraitis, Dmitry Toichkin, Adrien Journé, Yansong Chua, and Qinghai Guo · 2022
Later among the works it cites.
Scaling forward gradient with local losses
Mengye Ren, Simon Kornblith, Renjie Liao, and Geoffrey Hinton · 2022
Later among the works it cites.
Zeroth-order optimization with trajectory-informed derivative estimation
Yao Shu, Zhongxiang Dai, Weicong Sng, Arun Verma, Patrick Jaillet, and Bryan Kian Hsiang Low · 2022
Later among the works it cites.
Black-box tuning for language-model-as-a-service
Tianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang, and Xipeng Qiu · 2022
Later among the works it cites.
A comprehensive review of digital twin—part 1: modeling and twinning enabling technologies
Adam Thelen, Xiaoge Zhang, Olga Fink, Yan Lu, Sayan Ghosh, Byeng D Youn, Michael D Todd, Sankaran Mahadevan, Chao Hu, and Zhen Hu · 2022
Later among the works it cites.
Zeroth-order sciml: Non-intrusive integration of scientific software with deep learning
Ioannis Tsaknakis, Bhavya Kailkhura, Sijia Liu, Donald Loveland, James Diffenderfer, Anna Maria Hiszpanski, and Mingyi Hong · 2022
Later among the works it cites.
Zarts: On zero-order optimization for neural architecture search
Xiaoxing Wang, Wenxuan Guo, Jianlin Su, Xiaokang Yang, and Junchi Yan · 2022
Later among the works it cites.
Exploring parameter spaces with artificial intelligence and machine learning black-box optimization algorithms
Fernando Abreu de Souza, Miguel Crispim Romão, Nuno Filipe Castro, Mehraveh Nikjoo, and Werner Porod · 2023
Closest in time.
Compute-efficient deep learning: Algorithmic trends and opportunities
Brian R Bartoldson, Bhavya Kailkhura, and Davis Blalock · 2023
Closest in time.
Loss landscapes are all you need: Neural network generalization can be explained without the implicit bias of gradient descent
Ping-yeh Chiang, Renkun Ni, David Yu Miller, Arpit Bansal, Jonas Geiping, Micah Goldblum, and Tom Goldstein · 2023
Closest in time.
Fine-tuning language models with just forward passes
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora · 2023
Closest in time.
Certified zeroth-order black-box defense with robust unet denoiser
Astha Verma, Siddhesh Bangar, AV Subramanyam, Naman Lal, Rajiv Ratn Shah, and Shin’ichi Satoh · 2023
Closest in time.
Low-variance gradient estimation in unrolled computation graphs with es-single
Paul Vicol, Zico Kolter, and Kevin Swersky · 2023
Closest in time.
Revisiting zeroth-order optimization for memory-efficient llm fine-tuning: A benchmark
Yihua Zhang, Pingzhi Li, Junyuan Hong, Jiaxiang Li, Yimeng Zhang, Wenqing Zheng, Pin-Yu Chen, Jason D Lee, Wotao Yin, Mingyi Hong, et al · 2024
Closest in time.