Fetching the paper…
Reading the bibliography…
Backward propagation (BP) is widely used to compute the gradients in neural network training.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T. 1964 · 1964
Earlier work this paper cites.
Estimation of the mean of a multivariate normal distribution
Stein, C. M. 1981 · 1981
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E.; Hinton, G. E.; and Williams, R. J. 1986 · 1986
Earlier work this paper cites.
Explicit cost bounds of algorithms for multivariate tensor product problems
Wasilkowski, G. W.; and Wozniakowski, H. 1995 · 1995
Earlier work this paper cites.
Numerical integration using sparse grids
Gerstner, T.; and Griebel, M. 1998 · 1998
Earlier work this paper cites.
Artificial neural networks for solving ordinary and partial differential equations
Lagaris, I. E.; Likas, A.; and Fotiadis, D. I. 1998 · 1998
Earlier work this paper cites.
Picking winning tickets before training by preserving gradient flow
Wang, C.; Zhang, G.; and Grosse, R. 2020 · 2002
Earlier work this paper cites.
Sparse grid tutorial
Garcke, J.; et al. 2006 · 2006
Earlier work this paper cites.
A sparse grid stochastic collocation method for partial differential equations with random input data
Nobile, F.; Tempone, R.; and Webster, C. G. 2008 · 2008
Earlier work this paper cites.
Tensor-train decomposition
Oseledets, I. V. 2011 · 2011
Earlier work this paper cites.
Learning certified control using contraction metric
Sun, D.; Jha, S.; and Fan, C. 2020 · 2011
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S.; and Lan, G. 2013 · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I.; Martens, J.; Dahl, G.; and Hinton, G. 2013 · 2013
Earlier work this paper cites.
Decaf: A deep convolutional activation feature for generic visual recognition
Donahue, J.; Jia, Y.; Vinyals, O.; Hoffman, J.; Zhang, N.; Tzeng, E.; and Darrell, T. 2014 · 2014
Earlier work this paper cites.
Optimal rates for zero-order convex optimization: The power of two function evaluations
Duchi, J. C.; Jordan, M. I.; Wainwright, M. J.; and Wibisono, A. 2015 · 2015
Earlier work this paper cites.
A comprehensive linear speedup analysis for asynchronous stochastic parallel optimization from zeroth-order to first-order
Lian, X.; Zhang, H.; Hsieh, C.-J.; Huang, Y.; and Liu, J. 2016 · 2016
Earlier work this paper cites.
Random synaptic feedback weights support error backpropagation for deep learning
Lillicrap, T. P.; Cownden, D.; Tweed, D. B.; and Akerman, C. J. 2016 · 2016
Earlier work this paper cites.
Direct feedback alignment provides learning in deep neural networks
Nøkland, A. 2016 · 2016
Earlier work this paper cites.
ZOO: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models
Chen, P.-Y.; Zhang, H.; Sharma, Y.; Yi, J.; and Hsieh, C.-J. 2017 · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, B.; Moore, E.; Ramage, D.; Hampson, S.; and y Arcas, B. A. 2017 · 2017
Cited alongside, same era.
Random gradient-free minimization of convex functions
Nesterov, Y.; and Spokoiny, V. 2017 · 2017
Cited alongside, same era.
An optimal algorithm for bandit and zero-order convex optimization with two-point feedback
Shamir, O. 2017 · 2017
Cited alongside, same era.
Automatic differentiation in machine learning: a survey
Baydin, A. G.; Pearlmutter, B. A.; Radul, A. A.; and Siskind, J. M. 2018 · 2018
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems
Bernstein, J.; Wang, Y.-X.; Azizzadenesheli, K.; and Anandkumar, A. 2018 · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Zaken, E. B.; Ravfogel, S.; and Goldberg, Y. 2021 · 2021
Later among the works it cites.
On-FPGA training with ultra memory reduction: A low-precision tensor method
Zhang, K.; Hawkins, C.; Zhang, X.; Hao, C.; and Zhang, Z. 2021 · 2021
Later among the works it cites.
Zeroth-order nonconvex stochastic optimization: Handling constraints, high dimensionality, and saddle points
Balasubramanian, K.; and Ghadimi, S. 2022 · 2022
Later among the works it cites.
Gradients without backpropagation
Baydin, A. G.; Pearlmutter, B. A.; Syme, D.; Wood, F.; and Torr, P. 2022 · 2022
Later among the works it cites.
A theoretical and empirical comparison of gradient approximations in derivative-free optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Frankle, J.; and Carbin, M. 2018 · 2018
Cited alongside, same era.
Zo-adamm: Zeroth-order adaptive momentum method for black-box optimization
Chen, X.; Liu, S.; Xu, K.; Li, X.; Lin, X.; Hong, M.; and Cox, D. 2019 · 2019
Cited alongside, same era.
signSGD via zeroth-order oracle
Liu, S.; Chen, P.-Y.; Chen, X.; and Hong, M. 2019 · 2019
Cited alongside, same era.
Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations
Raissi, M.; Perdikaris, P.; and Karniadakis, G. E. 2019 · 2019
Cited alongside, same era.
A stochastic derivative-free optimization method with importance sampling: Theory and learning to control
Bibi, A.; Bergou, E. H.; Sener, O.; Ghanem, B.; and Richtarik, P. 2020 · 2020
Cited alongside, same era.
A mathematical model for automatic differentiation in machine learning
Bolte, J.; and Pauwels, E. 2020 · 2020
Cited alongside, same era.
FLOPS: EFficient On-Chip Learning for OPtical Neural Networks Through Stochastic Zeroth-Order Optimization
Gu, J.; Zhao, Z.; Feng, C.; Li, W.; Chen, R. T.; and Pan, D. Z. 2020 · 2020
Cited alongside, same era.
Berahas, A. S.; Cao, L.; Choromanski, K.; and Scheinberg, K. 2022 · 2022
Later among the works it cites.
CAN-PINN: A fast physics-informed neural network based on coupled-automatic–numerical differentiation method
Chiu, P.-H.; Wong, J. C.; Ooi, C.; Dao, M. H.; and Ong, Y.-S. 2022 · 2022
Later among the works it cites.
Towards compact neural networks via end-to-end training: A bayesian tensor approach with automatic rank determination
Hawkins, C.; Liu, X.; and Zhang, Z. 2022 · 2022
Later among the works it cites.
The forward-forward algorithm: Some preliminary investigations
Hinton, G. 2022 · 2022
Later among the works it cites.
Low dimensional trajectory hypothesis is true: Dnns can be trained in tiny subspaces
Li, T.; Tan, L.; Huang, Z.; Tao, Q.; Liu, Y.; and Huang, X. 2022 · 2022
Later among the works it cites.
Physics informed neural network using finite difference method
Lim, K. L.; Dutta, R.; and Rotaru, M. 2022 · 2022
Later among the works it cites.
On-device training under 256kb memory
Lin, J.; Zhu, L.; Chen, W.-M.; Wang, W.-C.; Gan, C.; and Han, S. 2022 · 2022
Later among the works it cites.
Scaling forward gradient with local losses
Ren, M.; Kornblith, S.; Liao, R.; and Hinton, G. 2022 · 2022
Later among the works it cites.
Xiang, Z.; Peng, W.; Zhou, W.; and Yao, W. 2022 · 2022
Later among the works it cites.
How to Robustify Black-Box ML Models? A Zeroth-Order Optimization Perspective
Zhang, Y.; Yao, Y.; Jia, J.; Yi, J.; Hong, M.; Chang, S.; and Liu, S. 2022 · 2022
Later among the works it cites.
Learning physics-informed neural networks without stacked back-propagation
He, D.; Li, S.; Shi, W.; Gao, X.; Zhang, J.; Bian, J.; Wang, L.; and Liu, T.-Y. 2023 · 2023
Closest in time.
Fine-Tuning Language Models with Just Forward Passes
Malladi, S.; Gao, T.; Nichani, E.; Damian, A.; Lee, J. D.; Chen, D.; and Arora, S. 2023 · 2023
Closest in time.
Vflh: A following-the-leader-history based algorithm for adaptive online convex optimization with stochastic constraints
Yang, Y.; Chen, L.; Zhou, P.; and Ding, X. 2023 · 2023
Closest in time.
PIFON-EPT: MR-Based Electrical Property Tomography Using Physics-Informed Fourier Networks
Yu, X.; Serrallés, J. E.; Giannakopoulos, I. I.; Liu, Z.; Daniel, L.; Lattanzi, R.; and Zhang, Z. 2023 · 2023
Closest in time.