Fetching the paper…
Reading the bibliography…
Using the notion of conservative gradient, we provide a simple model to estimate the computational costs of the backward and forward modes of algorithmic differentiation for a wide class of nonsmooth programs.
Programs for automatic differentiation for the machine BESM
L. M. Beda, L. N. Korolev, N. V. Sukkikh, and T. S. Frolova · 1959
Earlier work this paper cites.
A simple automatic derivative evaluation program
Robert Edwin Wengert · 1964
Earlier work this paper cites.
Gaussian elimination is not optimal
Volker Strassen et al · 1969
Earlier work this paper cites.
Checking the calculation of gradients
Philip Wolfe · 1982
Earlier work this paper cites.
The complexity of partial derivatives
Walter Baur and Volker Strassen · 1983
Earlier work this paper cites.
Optimization and nonsmooth analysis
Frank H Clarke · 1983
Earlier work this paper cites.
Learning Representations by Back-propagating Errors
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
On automatic differentiation
Andreas Griewank et al · 1989
Earlier work this paper cites.
On concepts of directional differentiability
Alexander Shapiro · 1990
Earlier work this paper cites.
Gradient calculations for dynamic recurrent neural networks: A survey
Barak A Pearlmutter · 1995
Earlier work this paper cites.
Differentiation of the cholesky algorithm
Stephen P Smith · 1995
Earlier work this paper cites.
The use of adjoint equations in numerical weather prediction
P. Courtier and F. Rabier · 1997
Earlier work this paper cites.
Theory of linear and integer programming
Alexander Schrijver · 1998
Earlier work this paper cites.
Piggyback differentiation and optimization
Andreas Griewank and Christèle Faure · 2003
Earlier work this paper cites.
Lexicographic differentiation of nonsmooth functions
Yu Nesterov · 2005
Earlier work this paper cites.
Toward an optimal algorithm for matrix multiplication
Sara Robinson · 2005
Earlier work this paper cites.
A review of the adjoint-state method for computing the gradient of a functional with geophysical applications
R-E Plessix · 2006
Earlier work this paper cites.
Evaluating derivatives: principles and techniques of algorithmic differentiation
Andreas Griewank and Andrea Walther · 2008
Earlier work this paper cites.
Evaluating an element of the clarke generalized jacobian of a piecewise differentiable function
Kamil A Khan and Paul I Barton · 2012
Earlier work this paper cites.
Introduction to piecewise differentiable equations
Stefan Scholtes · 2012
Cited alongside, same era.
Multiplying matrices faster than coppersmith-winograd
Virginia Vassilevska Williams · 2012
Cited alongside, same era.
Real algebraic geometry , volume 36
Jacek Bochnak, Michel Coste, and Marie-Françoise Roy · 2013
Cited alongside, same era.
Automated derivation of the adjoint of high-level transient finite element programs
Patrick E Farrell, David A Ham, Simon W Funke, and Marie E Rognes · 2013
Cited alongside, same era.
On stable piecewise linearization and generalized algorithmic differentiation
Andreas Griewank · 2013
Cited alongside, same era.
Evaluating an element of the clarke generalized jacobian of a composite piecewise differentiable function
Kamil A Khan and Paul I Barton · 2013
Cited alongside, same era.
Neural ordinary differential equations
Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud · 2018
Later among the works it cites.
Provably correct automatic sub-differentiation for qualified programs
Sham M Kakade and Jason D Lee · 2018
Later among the works it cites.
Branch-locking ad techniques for nonsmooth composite functions and nonsmooth implicit functions
Kamil A Khan · 2018
Later among the works it cites.
Differentiating through a cone program
Akshay Agrawal, Shane Barratt, Stephen Boyd, Enzo Busseti, and Walaa M Moursi · 2019
Later among the works it cites.
Deep equilibrium models
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2019
Later among the works it cites.
Treating artificial neural net training as a nonsmooth global optimization problem
A. Griewank and A. Rojas · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Differentiating the method of conjugate gradients
Serge Gratton, David Titley-Peloquin, Philippe Toint, and Jean Tshimanga Ilunga · 2014
Cited alongside, same era.
Powers of tensors and fast matrix multiplication
François Le Gall · 2014
Cited alongside, same era.
A vector forward mode of automatic differentiation for generalized derivative evaluation
Kamil A Khan and Paul I Barton · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Implicit differentiation of lasso-type models for hyperparameter optimization
Quentin Bertrand, Quentin Klopfenstein, Mathieu Blondel, Samuel Vaiter, Alexandre Gramfort, and Joseph Salmon · 2020
Later among the works it cites.
Beyond the oracle: Opportunities of piecewise differentiation
A. Griewank and A. Walther · 2020
Later among the works it cites.
Optimizing millions of hyperparameters by implicit differentiation
Jonathan Lorraine, Paul Vicol, and David Duvenaud · 2020
Later among the works it cites.
Automatic differentiation of some first-order methods in parametric optimization
Sheheryar Mehmood and Peter Ochs · 2020
Later among the works it cites.
Monotone operator equilibrium networks
Ezra Winston and J Zico Kolter · 2020
Later among the works it cites.
A refined laser method and faster matrix multiplication
Josh Alman and Virginia Vassilevska Williams · 2021
Later among the works it cites.
Numerical influence of relu’(0) on backpropagation
David Bertoin, Jérôme Bolte, Sébastien Gerchinovitz, and Edouard Pauwels · 2021
Later among the works it cites.
Efficient and modular implicit differentiation
Mathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig, Stephan Hoyer, Felipe Llinares-López, Fabian Pedregosa, and Jean-Philippe Vert · 2021
Later among the works it cites.
Nonsmooth implicit differentiation for machine-learning and optimization
Jérôme Bolte, Tam Le, Edouard Pauwels, and Tony Silveti-Falls · 2021
Later among the works it cites.
Conservative and semismooth derivatives are equivalent for semialgebraic maps
Damek Davis and Dmitriy Drusvyatskiy · 2021
Later among the works it cites.
The structure of conservative gradient fields
Adrian Lewis and Tonghua Tian · 2021
Later among the works it cites.
Guanhua Wang and Jeffrey A. Fessler · 2021
Later among the works it cites.