Fetching the paper…
Reading the bibliography…
Using backpropagation to compute gradients of objective functions for optimization has remained a mainstay of machine learning.
A simple automatic derivative evaluation program
Wengert, R. E · 1964
Earlier work this paper cites.
The representation of the cumulative rounding error of an algorithm as a Taylor expansion of the local rounding errors
Linnainmaa, S · 1970
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W · 1983
Earlier work this paper cites.
Optimization by simulated annealing
Kirkpatrick, S., Gelatt, C. D., and Vecchi, M. P · 1983
Earlier work this paper cites.
Learning internal representations by error propagation
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1985
Earlier work this paper cites.
Learning process in an asymmetric threshold network
LeCun, Y · 1986
Earlier work this paper cites.
Stochastic learning networks and their electronic implementation
Alspector, J., Allen, R., Hu, V., and Satyanarayana, S · 1987
Earlier work this paper cites.
PhD thesis: Modeles connexionnistes de l’apprentissage (connectionist learning models)
LeCun, Y · 1987
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Williams, R. J. and Zipser, D · 1989
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
Spall, J. C. et al · 1992
Earlier work this paper cites.
Fast exact multiplication by the Hessian
Pearlmutter, B. A · 1994
Earlier work this paper cites.
Gradient calculations for dynamic recurrent neural networks: A survey
Pearlmutter, B. A · 1995
Earlier work this paper cites.
The complex-step derivative approximation
Martins, J. R., Sturdza, P., and Alonso, J. J · 2003
Earlier work this paper cites.
What color is your Jacobian? Graph coloring for computing derivatives
Gebremedhin, A. H., Manne, F., and Pothen, A · 2005
Earlier work this paper cites.
Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation
Griewank, A. and Walther, A · 2008
Cited alongside, same era.
Who invented the reverse mode of differentiation?
Griewank, A · 2012
Cited alongside, same era.
Efficiency of coordinate descent methods on huge-scale optimization problems
Nesterov, Y · 2012
Cited alongside, same era.
How auto-encoders could provide credit assignment in deep networks via target propagation
Bengio, Y · 2014
Cited alongside, same era.
Adjoints by automatic differentiation
Hascoët, L · 2014
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Later among the works it cites.
SEGA: Variance reduction via gradient sketching
Hanzely, F., Mishchenko, K., and Richtárik, P · 2018
Later among the works it cites.
On first-order meta-learning algorithms
Nichol, A., Achiam, J., and Schulman, J · 2018
Later among the works it cites.
Divide-and-conquer checkpointing for arbitrary programs with no user annotation
Siskind, J. M. and Pearlmutter, B. A · 2018
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bengio, Y., Lee, D.-H., Bornschein, J., Mesnard, T., and Lin, Z · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Coordinate descent algorithms
Wright, S. J · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Cited alongside, same era.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Cited alongside, same era.
Random synaptic feedback weights support error backpropagation for deep learning
Lillicrap, T. P., Cownden, D., Tweed, D. B., and Akerman, C. J · 2016
Cited alongside, same era.
Decoupled neural interfaces using synthetic gradients
Jaderberg, M., Czarnecki, W. M., Osindero, S., Vinyals, O., Graves, A., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Deriving differential target propagation from iterating approximate inverses
Bengio, Y · 2020
Later among the works it cites.
Mathematics for Machine Learning
Deisenroth, M. P., Faisal, A. A., and Ong, C. S · 2020
Later among the works it cites.
Backpropagation and the brain
Lillicrap, T. P., Santoro, A., Marris, L., Akerman, C. J., and Hinton, G · 2020
Later among the works it cites.
A theoretical framework for target propagation
Meulemans, A., Carzaniga, F., Suykens, J., Sacramento, J., and Grewe, B. F · 2020
Later among the works it cites.
Monte Carlo gradient estimation in machine learning
Mohamed, S., Rosca, M., Figurnov, M., and Mnih, A · 2020
Later among the works it cites.
Who invented backpropagation?
Schmidhuber, J · 2020
Later among the works it cites.
Langevin Monte Carlo: random coordinate descent and variance reduction
Ding, Z. and Li, Q · 2021
Later among the works it cites.
An accelerated directional derivative method for smooth stochastic convex optimization
Dvurechensky, P., Gorbunov, E., and Gasnikov, A · 2021
Later among the works it cites.
Learning by directional gradient descent
Anonymous · 2022
Closest in time.