Fetching the paper…
Reading the bibliography…
Least squares is by far the simplest and most commonly applied computational method in many fields.
Methods of conjugate gradients for solving linear systems
M. Hestenes and E. Stiefel · 1952
Earlier work this paper cites.
On the numerical solution of heat conduction problems in two and three space variables
J. Douglas and H. Rachford · 1956
Earlier work this paper cites.
A simple automatic derivative evaluation program
Robert Edwin Wengert · 1964
Earlier work this paper cites.
Numerical methods for solving linear least squares problems
G. Golub · 1965
Earlier work this paper cites.
Brève communication. régularisation d’inéquations variationnelles par approximations successives
B. Martinet · 1970
Earlier work this paper cites.
On Bayesian methods for seeking the extremum
J. Močkus · 1975
Earlier work this paper cites.
Solution of sparse indefinite systems of linear equations
C. Paige and M. Saunders · 1975
Earlier work this paper cites.
Splitting algorithms for the sum of two nonlinear operators
P. Lions and B. Mercier · 1979
Earlier work this paper cites.
A direct method for the solution of sparse linear least squares problems
A. Björck and I. Duff · 1980
Earlier work this paper cites.
Compiling fast partial derivatives of functions given by algorithms
B. Speelpenning · 1980
Earlier work this paper cites.
Lsqr: An algorithm for sparse linear equations and sparse least squares
C. Paige and M. Saunders · 1982
Earlier work this paper cites.
The complexity of partial derivatives
W. Baur and V. Strassen · 1983
Earlier work this paper cites.
Minimization methods for non-differentiable functions
N. Shor · 1985
Earlier work this paper cites.
Introduction to optimization
B. Polyak · 1987
Earlier work this paper cites.
Algorithm 679: A set of level 3 basic linear algebra subprograms: model implementation and test programs
J. Dongarra, J. Cruz, S. Hammarling, and I. Duff · 1990
Earlier work this paper cites.
Solving least squares problems
C. Lawson and R. Hanson · 1995
Earlier work this paper cites.
Differentiation of the cholesky algorithm
S. Smith · 1995
Earlier work this paper cites.
Adapting arbitrary normal mutation distributions in evolution strategies: The covariance matrix adaptation
N. Hansen and A. Ostermeier · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Adaptive regularization in neural network modeling
J. Larsen, C. Svarer, L. Andersen, and L. Hansen · 1998
Earlier work this paper cites.
Early stopping-but when?
L. Prechelt · 1998
Earlier work this paper cites.
LAPACK Users’ guide
E. Anderson, Z. Bai, C. Bischof, L. Blackford, J. Demmel, J. Dongarra, J. Du Croz, A. Greenbaum, S. Hammarling, and A. McKenney · 1999
Earlier work this paper cites.
Gradient based adaptive regularization
R. Eigenmann and J. Nossek · 1999
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Y. Bengio · 2000
Earlier work this paper cites.
Choosing multiple parameters for support vector machines
O. Chapelle, V. Vapnik, O. Bousquet, and S. Mukherjee · 2002
Cited alongside, same era.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Cited alongside, same era.
Gaussian Processes in Machine Learning
C. Rasmussen · 2004
Cited alongside, same era.
Numerical optimization
J. Nocedal and S. Wright · 2006
Cited alongside, same era.
An efficient method for gradient-based adaptation of hyperparameters in SVM models
S. Keerthi, V. Sindhwani, and O. Chapelle · 2007
Cited alongside, same era.
Efficient multiple hyperparameter learning for log-linear models
C. Foo, C. Do, and A. Ng · 2008
Cited alongside, same era.
Structured prediction energy networks
D. Belanger and A. McCallum · 2016
Later among the works it cites.
J. Fu, H. Luo, J. Feng, and T. Chua · 2016
Later among the works it cites.
Scalable gradient-based tuning of continuous regularization hyperparameters
J. Luketina, M. Berglund, K. Greff, and T. Raiko · 2016
Later among the works it cites.
Hyperparameter optimization with approximate gradient
F. Pedregosa · 2016
Later among the works it cites.
Optnet: Differentiable optimization as a layer in neural networks
B. Amos and Z. Kolter · 2017
Later among the works it cites.
A fast and differentiable qp solver for pytorch
B. Amos · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Griewank and A. Walther · 2008
Cited alongside, same era.
Implicit functions and solution mappings
A. Dontchev and T. Rockafellar · 2009
Cited alongside, same era.
CUDA by example: an introduction to general-purpose GPU programming
J. Sanders and E. Kandrot · 2010
Cited alongside, same era.
Distributed optimization and statistical learning via the alternating direction method of multipliers
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein · 2011
Cited alongside, same era.
The numpy array: a structure for efficient numerical computation
S. Walt, S. Colbert, and G. Varoquaux · 2011
Cited alongside, same era.
Random search for hyper-parameter optimization
J. Bergstra and Y. Bengio · 2012
Cited alongside, same era.
Later among the works it cites.
End-to-end learning for structured prediction energy networks
D. Belanger, B. Yang, and A. McCallum · 2017
Later among the works it cites.
Task-based end-to-end model learning in stochastic optimization
P. Donti, B. Amos, and Z. Kolter · 2017
Later among the works it cites.
Lecture 11 notes for ee104, 2017
S. Lall and S. Boyd · 2017
Later among the works it cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Later among the works it cites.
Differentiable mpc for end-to-end planning and control
B. Amos, I. Jimenez, J. Sacks, B. Boots, and Z. Kolter · 2018
Later among the works it cites.
On the differentiability of the solution to convex optimization problems
S. Barratt · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. Curtis, and J. Nocedal · 2018
Later among the works it cites.
Automatic differentiation in machine learning: a survey
A. Baydin, B. Pearlmutter, A. Radul, and J. Siskind · 2018
Later among the works it cites.
Introduction to applied linear algebra: vectors, matrices, and least squares
S. Boyd and L. Vandenberghe · 2018
Later among the works it cites.
End-to-end differentiable physics for learning and control
F. de Avila Belbute-Peres, K. Smith, K. Allen, J. Tenenbaum, and Z. Kolter · 2018
Later among the works it cites.
Don’t unroll adjoint: Differentiating ssa-form programs
M. Innes · 2018
Later among the works it cites.
Stochastic hyperparameter optimization through hypernetworks
J. Lorraine and D. Duvenaud · 2018
Later among the works it cites.
What game are we playing? end-to-end learning in normal and extensive form games
C. Ling, F. Fang, and Z. Kolter · 2018
Later among the works it cites.
Learning to reweight examples for robust deep learning
M. Ren, W. Zeng, B. Yang, and R. Urtasun · 2018
Later among the works it cites.
Tangent: Automatic differentiation using source-code transformation for dynamically typed array programming
B. van Merriënboer, D. Moldovan, and A. Wiltschko · 2018
Later among the works it cites.
Tensorflow eager: A multi-stage, python-embedded dsl for machine learning
A. Agrawal, A. Modi, A. Passos, A. Lavoie, A. Agarwal, A. Shankar, A. Ganichev, J. Levenberg, M. Hong, R. Monga, and S. Cai · 2019
Closest in time.
Large scale learning of agent rationality in two-player zero-sum games
C. Ling, F. Fang, and Z. Kolter · 2019
Closest in time.