Fetching the paper…
Reading the bibliography…
Automatic differentiation (autodiff) has revolutionized machine learning.
Methods of conjugate gradients for solving linear systems
M. R. Hestenes, E. Stiefel, et al · 1952
Earlier work this paper cites.
A simple automatic derivative evaluation program
R. E. Wengert · 1964
Earlier work this paper cites.
Proximité et dualité dans un espace hilbertien
J.-J. Moreau · 1965
Earlier work this paper cites.
Optimization and Nonsmooth Analysis
F. Clarke · 1983
Earlier work this paper cites.
An O ( n ) O(n) algorithm for quadratic knapsack problems
P. Brucker · 1984
Earlier work this paper cites.
Projections onto order simplexes
S. Grotzinger and C. Witzgall · 1984
Earlier work this paper cites.
A finite algorithm for finding the projection of a point onto the canonical simplex of ℝ n \mathbb{R}^{n}
C. Michelot · 1986
Earlier work this paper cites.
Gmres: A generalized minimal residual algorithm for solving nonsymmetric linear systems
Y. Saad and M. H. Schultz · 1986
Earlier work this paper cites.
Bi-CGSTAB: A fast and smoothly converging variant of Bi-CG for the solution of nonsymmetric linear systems
H. A. v. d. Vorst and H. A. van der Vorst · 1992
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Y. Bengio · 2000
Earlier work this paper cites.
Minimizing separable convex functions subject to simple chain constraints
M. J. Best, N. Chakravarti, and V. A. Ubhaya · 2000
Earlier work this paper cites.
On the algorithmic implementation of multiclass kernel-based vector machines
K. Crammer and Y. Singer · 2001
Earlier work this paper cites.
Choosing multiple parameters for support vector machines
O. Chapelle, V. Vapnik, O. Bousquet, and S. Mukherjee · 2002
Earlier work this paper cites.
Accuracy and Stability of Numerical Algorithms
N. J. Higham · 2002
Earlier work this paper cites.
Design of experiments of the nips 2003 variable selection benchmark
I. Guyon · 2003
Earlier work this paper cites.
Least angle regression
B. Efron, T. Hastie, I. Johnstone, and R. Tibshirani · 2004
Earlier work this paper cites.
Structural relaxation made simple
E. Bitzek, P. Koskinen, F. Gähler, M. Moseler, and P. Gumbsch · 2006
Earlier work this paper cites.
Algorithmic differentiation of implicit functions and optimal values
B. M. Bell and J. V. Burke · 2008
Earlier work this paper cites.
Efficient projections onto the ℓ 1 \ell_{1} -ball for learning in high dimensions
J. C. Duchi, S. Shalev-Shwartz, Y. Singer, and T. Chandra · 2008
Earlier work this paper cites.
Liblinear: A library for large linear classification
R.-E. Fan, K.-W. Chang, C.-J. Hsieh, X.-R. Wang, and C.-J. Lin · 2008
Earlier work this paper cites.
Evaluating derivatives: principles and techniques of algorithmic differentiation
A. Griewank and A. Walther · 2008
Earlier work this paper cites.
Cross-validation optimization for large scale structured classification kernel methods
M. W. Seeger · 2008
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Generic methods for optimization-based modeling
J. Domke · 2012
Earlier work this paper cites.
The implicit function theorem: history, theory, and applications
S. G. Krantz and H. R. Parks · 2012
Earlier work this paper cites.
Task-driven dictionary learning
J. Mairal, F. Bach, and J. Ponce · 2012
Earlier work this paper cites.
Complexity analysis of the lasso regularization path
J. Mairal and B. Yu · 2012
Earlier work this paper cites.
Perturbation analysis of optimization problems
J. F. Bonnans and A. Shapiro · 2013
Earlier work this paper cites.
Sinkhorn distances: lightspeed computation of optimal transport
M. Cuturi · 2013
Earlier work this paper cites.
The lasso problem and uniqueness
R. J. Tibshirani · 2013
Cited alongside, same era.
Local behavior of sparse analysis regularization: Applications to risk estimation
S. Vaiter, C.-A. Deledalle, G. Peyré, C. Dossal, and J. Fadili · 2013
Cited alongside, same era.
Stein unbiased gradient estimator of the risk (sugar) for multiple parameter selection
C.-A. Deledalle, S. Vaiter, J. Fadili, and G. Peyré · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Proximal algorithms
N. Parikh and S. Boyd · 2014
Cited alongside, same era.
Supervised non-euclidean sparse nmf via bilevel optimization with applications to speech enhancement
P. Sprechmann, A. M. Bronstein, and G. Sapiro · 2014
Differentiating through a cone program
A. Agrawal, S. Barratt, S. Boyd, E. Busseti, and W. M. Moursi · 2019
Later among the works it cites.
Differentiable optimization-based modeling for machine learning
B. Amos · 2019
Later among the works it cites.
Casadi: a software framework for nonlinear optimization and optimal control
J. A. Andersson, J. Gillis, G. Horn, J. B. Rawlings, and M. Diehl · 2019
Later among the works it cites.
S. Bai, J. Z. Kolter, and V. Koltun · 2019
Later among the works it cites.
Structured prediction with projection oracles
M. Blondel · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fast projection onto the simplex and the ℓ 1 \ell_{1} ball
L. Condat · 2016
Cited alongside, same era.
Cvxpy: A python-embedded modeling language for convex optimization
S. Diamond and S. Boyd · 2016
Cited alongside, same era.
S. Gould, B. Fernando, A. Cherian, P. Anderson, R. S. Cruz, and E. Guo · 2016
Cited alongside, same era.
Efficient bregman projections onto the permutahedron and related polytopes
C. H. Lim and S. J. Wright · 2016
Cited alongside, same era.
From softmax to sparsemax: A sparse model of attention and multi-label classification
A. F. Martins and R. F. Astudillo · 2016
Cited alongside, same era.
Conic optimization via operator splitting and homogeneous self-dual embedding
B. O’Donoghue, E. Chu, N. Parikh, and S. Boyd · 2016
Cited alongside, same era.
L. El Ghaoui, F. Gu, B. Travacca, A. Askari, and A. Y. Tsai · 2019
Later among the works it cites.
Deep declarative networks: A new hope
S. Gould, R. Hartley, and D. Campbell · 2019
Later among the works it cites.
Neural reparameterization improves structural optimization
S. Hoyer, J. Sohl-Dickstein, and S. Greydanus · 2019
Later among the works it cites.
Meta-learning with implicit gradients
A. Rajeswaran, C. Finn, S. Kakade, and S. Levine · 2019
Later among the works it cites.
Deep network classification by scattering and homotopy dictionary learning
J. Zarka, L. Thiry, T. Angles, and S. Mallat · 2019
Later among the works it cites.
Super-efficiency of automatic differentiation for functions defined as a minimum
P. Ablin, G. Peyré, and T. Moreau · 2020
Later among the works it cites.
Learning composable energy surrogates for pde order reduction
A. Beatson, J. Ash, G. Roeder, T. Xue, and R. P. Adams · 2020
Later among the works it cites.
Implicit differentiation of lasso-type models for hyperparameter optimization
Q. Bertrand, Q. Klopfenstein, M. Blondel, S. Vaiter, A. Gramfort, and J. Salmon · 2020
Later among the works it cites.
Fast differentiable sorting and ranking
M. Blondel, O. Teboul, Q. Berthet, and J. Djolonga · 2020
Later among the works it cites.
Learning to solve tv regularised problems with unrolled algorithms
H. Cherkaoui, J. Sulam, and T. Moreau · 2020
Later among the works it cites.
Deep implicit layers tutorial - neural ODEs, deep equilibirum models, and beyond
D. Duvenaud, J. Z. Kolter, and M. Johnson · 2020
Later among the works it cites.
On the iteration complexity of hypergradient computation
R. Grazzi, L. Franceschi, M. Pontil, and S. Salzo · 2020
Later among the works it cites.
Optimizing millions of hyperparameters by implicit differentiation
J. Lorraine, P. Vicol, and D. Duvenaud · 2020
Later among the works it cites.
Lp-sparsemap: Differentiable relaxed optimization for sparse structured prediction
V. Niculae and A. Martins · 2020
Later among the works it cites.
Jax md: A framework for differentiable physics
S. Schoenholz and E. D. Cubuk · 2020
Later among the works it cites.
Implicit differentiation for fast hyperparameter selection in non-smooth convex learning
Q. Bertrand, Q. Klopfenstein, M. Massias, M. Blondel, S. Vaiter, A. Gramfort, and J. Salmon · 2021
Closest in time.
Nonsmooth implicit differentiation for machine-learning and optimization
J. Bolte, T. Le, E. Pauwels, and T. Silveti-Falls · 2021
Closest in time.
Variational data assimilation with a learned inverse observation operator
T. Frerix, D. Kochkov, J. A. Smith, D. Cremers, M. P. Brenner, and S. Hoyer · 2021
Closest in time.
Decomposing reverse-mode automatic differentiation
R. Frostig, M. Johnson, D. Maclaurin, A. Paszke, and A. Radul · 2021
Closest in time.
Fixed point networks: Implicit depth models with jacobian-free backprop
S. W. Fung, H. Heaton, Q. Li, D. McKenzie, S. Osher, and W. Yin · 2021
Closest in time.
On training implicit models
Z. Geng, X.-Y. Zhang, S. Bai, Y. Wang, and Z. Lin · 2021
Closest in time.
Bilevel optimization: Convergence analysis and enhanced design
K. Ji, J. Yang, and Y. Liang · 2021
Closest in time.
Z. Ramzi, F. Mannel, S. Bai, J.-L. Starck, P. Ciuciu, and T. Moreau · 2021
Closest in time.