Fetching the paper…
Reading the bibliography…
Artificial intelligence has recently experienced remarkable advances, fueled by large models, vast datasets, accelerated hardware, and, last but not least, the transformative power of differentiable programming.
“Variabilità e mutabilità”
Corrado Gini · 1912
Earlier work this paper cites.
“Über partielle und totale differenzierbarkeit von Funktionen mehrerer Variabeln und über die Transformation der Doppelintegrale”
Hans Rademacher · 1919
Earlier work this paper cites.
“Generalized poisson distribution”
FE Satterthwaite · 1942
Earlier work this paper cites.
“A method for the solution of certain non-linear problems in least squares”
Kenneth Levenberg · 1944
Earlier work this paper cites.
“Rings of real-valued continuous functions. I”
Edwin Hewitt · 1948
Earlier work this paper cites.
“A mathematical theory of communication”
Claude Shannon · 1948
Earlier work this paper cites.
“Methods of conjugate gradients for solving linear systems”
Magnus Hestenes and Eduard Stiefel · 1952
Earlier work this paper cites.
“The formula for change in variables in a multiple integral”
J Schwartz · 1954
Earlier work this paper cites.
“A short account of the history of mathematics”
Walter Ball · 1960
Earlier work this paper cites.
“Uncertainty, information, and sequential experiments”
Morris DeGroot · 1962
Earlier work this paper cites.
“An algorithm for least-squares estimation of nonlinear parameters”
Donald Marquardt · 1963
Earlier work this paper cites.
“Gradient methods for the minimisation of functionals”
Boris Polyak · 1963
Earlier work this paper cites.
“Some methods of speeding up the convergence of iteration methods”
Boris Polyak · 1964
Earlier work this paper cites.
“Statistical inference for probabilistic functions of finite state Markov chains”
Leonard Baum and Ted Petrie · 1966
Earlier work this paper cites.
“Error bounds for convolutional codes and an asymptotically optimum decoding algorithm”
Andrew Viterbi · 1967
Earlier work this paper cites.
“The convergence of a class of double-rank minimization algorithms 1. general considerations”
Charles Broyden · 1970
Earlier work this paper cites.
“A new approach to variable metric algorithms”
Roger Fletcher · 1970
Earlier work this paper cites.
“A family of variable-metric methods derived by variational means”
Donald Goldfarb · 1970
Earlier work this paper cites.
“Conditioning of quasi-Newton methods for function minimization”
David Shanno · 1970
Earlier work this paper cites.
“Differentiation under the integral sign”
Harley Flanders · 1973
Earlier work this paper cites.
“The viterbi algorithm”
G Forney · 1973
Earlier work this paper cites.
“Generalized gradients and applications”
Frank Clarke · 1975
Earlier work this paper cites.
“Gaussian processes”
Takeyuki Hida and Masuyuki Hitsuda · 1976
Earlier work this paper cites.
“Orthogonal recurrent neural networks with scaled Cayley transform”
Kyle Helfrich, Devin Willmott and Qiang Ye · 1978
Earlier work this paper cites.
“Introduction to numerical analysis”
Josef Stoer et al · 1980
Earlier work this paper cites.
“The complexity of partial derivatives”
Walter Baur and Volker Strassen · 1983
Earlier work this paper cites.
“Differential equations and their applications”
Martin Braun and Martin Golubitsky · 1983
Earlier work this paper cites.
“Problem complexity and method efficiency in optimization”
Arkadi Nemirovski and David Yudin · 1983
Earlier work this paper cites.
“An O ( n ) O(n) algorithm for quadratic knapsack problems”
Peter Brucker · 1984
Earlier work this paper cites.
“How to compute fast a function and all its derivatives: A variation on the theorem of Baur-Strassen”
Jacques Morgenstern · 1985
Earlier work this paper cites.
“The mathematical theory of optimal processes and differential games”
Lev Pontryagin · 1985
Earlier work this paper cites.
“Conception optimale ou identification de formes, calcul rapide de la dérivée directionnelle de la fonction coût”
Jean Céa · 1986
Earlier work this paper cites.
“A finite algorithm for finding the projection of a point onto the canonical simplex of ℝ n \mathbb{R}^{n} ”
Christian Michelot · 1986
Earlier work this paper cites.
“GMRES: A generalized minimal residual algorithm for solving nonsymmetric linear systems”
Youcef Saad and Martin Schultz · 1986
Earlier work this paper cites.
“Abstract dynamic programming models under commutativity conditions”
Sergio Verdu and H Poor · 1987
Earlier work this paper cites.
“Improving the convergence of back-propagation learning with second order methods”
Sue Becker and Yann Le · 1988
Earlier work this paper cites.
“Antithetic acceleration of Monte Carlo integration in Bayesian inference”
John Geweke · 1988
Earlier work this paper cites.
“A theoretical framework for back-propagation”
Yann LeCun · 1988
Earlier work this paper cites.
“Possible generalization of Boltzmann-Gibbs statistics”
Constantino Tsallis · 1988
Earlier work this paper cites.
“Scans as primitive parallel operations”
Guy Blelloch · 1989
Earlier work this paper cites.
“A fast ‘Monte-Carlo cross-validation’procedure for large least squares problems with noisy data”
A Girard · 1989
Earlier work this paper cites.
“Exact maximum a posteriori estimation for binary images”
Dorothy Greig, Bruce Porteous and Allan Seheult · 1989
Earlier work this paper cites.
“A stochastic estimator of the trace of the influence matrix for Laplacian smoothing splines”
Michael Hutchinson · 1989
Earlier work this paper cites.
“On the limited memory method for large scale optimization”
Dong. Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
“A tutorial on hidden Markov models and selected applications in speech recognition”
Lawrence Rabiner · 1989
Earlier work this paper cites.
“Backpropagation through time: what it does and how to do it”
Paul Werbos · 1990
Earlier work this paper cites.
“Achieving logarithmic growth of temporal and spatial complexity in reverse automatic differentiation”
Andreas Griewank · 1992
Earlier work this paper cites.
“Bi-CGSTAB: A Fast and Smoothly Converging Variant of Bi-CG for the Solution of Nonsymmetric Linear Systems”
H Vorst and H van Vorst · 1992
Earlier work this paper cites.
“A history of mathematical notations”
Florian Cajori · 1993
Earlier work this paper cites.
“Convex analysis and minimization algorithms II”
Jean-Baptiste Hiriart-Urruty and Claude Lemaréchal · 1993
Earlier work this paper cites.
“The roots of backpropagation: from ordered derivatives to neural networks and political forecasting”
Paul Werbos · 1994
Earlier work this paper cites.
“Iterative methods for linear and nonlinear equations”
Carl Kelley · 1995
Earlier work this paper cites.
“Fuzzy sets and fuzzy logic”
George Klir and Bo Yuan · 1995
Earlier work this paper cites.
“Optimal time and minimum space-time product for reversing a certain class of programs”, 1996
José Grimm, Loı̈c Pottier and Nicole Rostaing-Schmidt · 1996
Earlier work this paper cites.
“Regression shrinkage and selection via the lasso”
Robert Tibshirani · 1996
Earlier work this paper cites.
“Factor graphs and algorithms”
Brendan Frey, Frank Kschischang, Hans-Andrea Loeliger and Niclas Wiberg · 1997
Earlier work this paper cites.
“Faster than the fast Legendre transform, the linear-time Legendre transform”
Yves Lucet · 1997
Earlier work this paper cites.
“Natural gradient works efficiently in learning”
Shun-Ichi Amari · 1998
Earlier work this paper cites.
“Using complex variables to estimate derivatives of real functions”
William Squire and George Trapp · 1998
Earlier work this paper cites.
“Policy gradient methods for reinforcement learning with function approximation”
Richard Sutton, David McAllester, Satinder Singh and Yishay Mansour · 1999
Earlier work this paper cites.
“Numerical optimization”
Stephen Wright and Jorge Nocedal · 1999
Earlier work this paper cites.
“The generalized distributive law”
Srinivas Aji and Robert McEliece · 2000
Earlier work this paper cites.
“Conditional random fields: Probabilistic models for segmenting and labeling sequence data”, 2001
John Lafferty, Andrew McCallum and Fernando Pereira · 2001
Earlier work this paper cites.
“Sharp uniform convexity and smoothness inequalities for trace norms”
Keith Ball, Eric Carlen and Elliott Lieb · 2002
Earlier work this paper cites.
“Learning with kernels: support vector machines, regularization, optimization, and beyond”
Bernhard Schölkopf and Alexander Smola · 2002
Earlier work this paper cites.
“Differential forms and the change of variable formula for multiple integrals”
Michael Taylor · 2002
Earlier work this paper cites.
“A mathematical view of automatic differentiation”
Andreas Griewank · 2003
Earlier work this paper cites.
“The complex-step derivative approximation”
Joaquim Martins, Peter Sturdza and Juan Alonso · 2003
Cited alongside, same era.
“Convex optimization”
Stephen Boyd and Lieven Vandenberghe · 2004
Cited alongside, same era.
“Game theory, maximum entropy, minimum discrepancy and robust Bayesian decision theory”
Peter Grünwald and A Dawid · 2004
Cited alongside, same era.
“An introduction to factor graphs”
H-A Loeliger · 2004
Cited alongside, same era.
“Kernel methods for pattern analysis”
John Shawe-Taylor and Nello Cristianini · 2004
Cited alongside, same era.
“Smooth minimization of non-smooth functions”
Yu Nesterov · 2005
Cited alongside, same era.
“Model selection and estimation in regression with grouped variables”
“The reversible residual network: Backpropagation without storing activations”
Aidan Gomez, Mengye Ren, Raquel Urtasun and Roger Grosse · 2017
Later among the works it cites.
“ Software 2.0 ”, 2017
Andrej Karpathy · 2017
Later among the works it cites.
“Random gradient-free minimization of convex functions”
Yurii Nesterov and Vladimir Spokoiny · 2017
Later among the works it cites.
“Nesterov’s Punctuated Equilibrium”, 2017
Benjamin Recht and Roy Frostig · 2017
Later among the works it cites.
“Evolution strategies as a scalable alternative to reinforcement learning”
Tim Salimans et al · 2017
Later among the works it cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ming Yuan and Yi Lin · 2006
Cited alongside, same era.
“An estimator for the diagonal of a matrix”
Costas Bekas, Effrosyni Kokiopoulou and Yousef Saad · 2007
Cited alongside, same era.
“Modified Gauss–Newton scheme with worst case guarantees for global performance”
Yu Nesterov · 2007
Cited alongside, same era.
“Nonsmooth analysis and control theory”
Francis Clarke, Yuri Ledyaev, Ronald Stern and Peter Wolenski · 2008
Cited alongside, same era.
“ Efficient projections onto the ℓ 1 \ell_{1} -ball for learning in high dimensions ”
John Duchi, Shai Shalev-Shwartz, Yoram Singer and Tushar Chandra · 2008
Cited alongside, same era.
“An introduction to functional derivatives”
Béla Frigyik, Santosh Srivastava and Maya Gupta · 2008
Cited alongside, same era.
“Automatic differentiation in machine learning: a survey”
Atilim Baydin, Barak Pearlmutter, Alexey Radul and Jeffrey Siskind · 2018
Later among the works it cites.
“JAX: composable transformations of Python+NumPy programs”, 2018
James Bradbury et al · 2018
Later among the works it cites.
“Neural ordinary differential equations”
Ricky Chen, Yulia Rubanova, Jesse Bettencourt and David Duvenaud · 2018
Later among the works it cites.
“Dice: The infinitely differentiable monte carlo estimator”
Jakob Foerster et al · 2018
Later among the works it cites.
“ Deep Learning est mort. Vive Differentiable Programming! ”, 2018
Yann LeCun · 2018
Later among the works it cites.
“An introduction to probabilistic programming”
Jan-Willem van Meent, Brooks Paige, Hongseok Yang and Frank Wood · 2018
Later among the works it cites.
“Differentiable dynamic programming for structured prediction and attention”
Arthur Mensch and Mathieu Blondel · 2018
Later among the works it cites.
“Lectures on convex optimization”
Yurii Nesterov · 2018
Later among the works it cites.
“On the fenchel duality between strong convexity and lipschitz continuou s gradient”
Xingyu Zhou · 2018
Later among the works it cites.
“Structured prediction with projection oracles”
Mathieu Blondel · 2019
Later among the works it cites.
“Backpack: Packing more into backprop”
Felix Dangel, Frederik Kunstner and Philipp Hennig · 2019
Later among the works it cites.
“Efficiency of minimizing compositions of convex functions and smooth maps”
Dmitriy Drusvyatskiy and Courtney Paquette · 2019
Later among the works it cites.
“ANODE: Unconditionally Accurate Memory-Efficient Gradients for Neural ODEs”
Amir Gholaminejad, Kurt Keutzer and George Biros · 2019
Later among the works it cites.
“Normalizing flows: Introduction and ideas”
Ivan Kobyzev, Simon Prince and Marcus Brubaker · 2019
Later among the works it cites.
“Limitations of the empirical Fisher approximation for natural gradient descent”
Frederik Kunstner, Philipp Hennig and Lukas Balles · 2019
Later among the works it cites.
“PyTorch: An Imperative Style, High-Performance Deep Learning Library”
Adam Paszke et al · 2019
Later among the works it cites.
“Computational optimal transport: With applications to data science”
Gabriel Peyré and Marco Cuturi · 2019
Later among the works it cites.
“Painless stochastic gradient: Interpolation, line-search, and convergence rates”
Sharan Vaswani et al · 2019
Later among the works it cites.
“Fine-tuning language models from human preferences”
Daniel Ziegler et al · 2019
Later among the works it cites.
“Learning with differentiable pertubed optimizers”
Quentin Berthet et al · 2020
Later among the works it cites.
“Learning with fenchel-young losses”
Mathieu Blondel, André Martins and Vlad Niculae · 2020
Later among the works it cites.
“A mathematical model for automatic differentiation in machine learning”
Jérôme Bolte and Edouard Pauwels · 2020
Later among the works it cites.
“Time dependence in non-autonomous neural odes”
Jared Davis et al · 2020
Later among the works it cites.
“Mathematics for machine learning”
Marc Deisenroth, A Faisal and Cheng Ong · 2020
Later among the works it cites.
“New insights and perspectives on the natural gradient method”
James Martens · 2020
Later among the works it cites.
“Monte carlo gradient estimation in machine learning”
Shakir Mohamed, Mihaela Rosca, Michael Figurnov and Andriy Mnih · 2020
Later among the works it cites.
“Gradient estimation with stochastic softmax tricks”
Max Paulus et al · 2020
Later among the works it cites.
“Mathematical foundations of data sciences”
Gabriel Peyré · 2020
Later among the works it cites.
“SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python”
Pauli Virtanen et al · 2020
Later among the works it cites.
“The implicit and explicit regularization effects of dropout”
Colin Wei, Sham Kakade and Tengyu Ma · 2020
Later among the works it cites.
“Second-order optimization for non-convex machine learning: An empirical study”
Peng Xu, Fred Roosta and Michael Mahoney · 2020
Later among the works it cites.
“Efficient and Modular Implicit Differentiation”
Mathieu Blondel et al · 2021
Later among the works it cites.
“Decomposing reverse-mode automatic differentiation”
Roy Frostig et al · 2021
Later among the works it cites.
“Storchastic: A framework for general stochastic automatic differentiation”
Emile Krieken, Jakub Tomczak and Annette Ten · 2021
Later among the works it cites.
“Survey of sequential convex programming and generalized Gauss-Newton methods”
Florian Messerer, Katrin Baumgärtner and Moritz Diehl · 2021
Later among the works it cites.
“Hutch++: Optimal stochastic trace estimation”
Raphael Meyer, Cameron Musco, Christopher Musco and David Woodruff · 2021
Later among the works it cites.
“Normalizing flows for probabilistic modeling and inference”
George Papamakarios et al · 2021
Later among the works it cites.
“Learning with algorithmic supervision via continuous relaxations”
Felix Petersen, Christian Borgelt, Hilde Kuehne and Oliver Deussen · 2021
Later among the works it cites.
“Anderson acceleration for contractive and noncontractive operators”
Sara Pollock and Leo Rebholz · 2021
Later among the works it cites.
“Momentum residual neural networks”
Michael Sander, Pierre Ablin, Mathieu Blondel and Gabriel Peyré · 2021
Later among the works it cites.
“Momentum residual neural networks”
Michael Sander, Pierre Ablin, Mathieu Blondel and Gabriel Peyré · 2021
Later among the works it cites.
“Linear transformers are secretly fast weight programmers”
Imanol Schlag, Kazuki Irie and Jürgen Schmidhuber · 2021
Later among the works it cites.
“Unbiased gradient estimation in unrolled computation graphs with persistent evolution strategies”
Paul Vicol, Luke Metz and Jascha Sohl-Dickstein · 2021
Later among the works it cites.
Aston Zhang, Zachary Lipton, Mu Li and Alexander Smola · 2021
Later among the works it cites.
“Mali: A memory efficient and reverse accurate integrator for neural odes”
Juntang Zhuang, Nicha Dvornek, Sekhar Tatikonda and James Duncan · 2021
Later among the works it cites.
“Stochastic diagonal estimation: probabilistic bounds and an improved algorithm”
Robert Baston and Yuji Nakatsukasa · 2022
Later among the works it cites.
“Gradients without backpropagation”
Atılımüneş Baydin et al · 2022
Later among the works it cites.
“On the complexity of nonsmooth automatic differentiation”
Jérôme Bolte, Ryan Boustany, Edouard Pauwels and Béatrice Pesquet-Popescu · 2022
Later among the works it cites.
“HesScale: Scalable Computation of Hessian Diagonals”
Mohamed Elsayed and A Mahmood · 2022
Later among the works it cites.
“Probabilistic Machine Learning: An introduction”
Kevin. Murphy · 2022
Later among the works it cites.
“You only linearize once: Tangents transpose to gradients”
Alexey Radul et al · 2022
Later among the works it cites.
“Differentiable programming à la Moreau”
Vincent Roulet and Zaid Harchaoui · 2022
Later among the works it cites.
“An introduction to optimization on smooth manifolds”
Nicolas Boumal · 2023
Later among the works it cites.
“Scaling vision transformers to 22 billion parameters”
Mostafa Dehghani et al · 2023
Later among the works it cites.
“XTrace: Making the most of every sample in stochastic trace estimation”
Ethan Epperly, Joel Tropp and Robert Webber · 2023
Later among the works it cites.
“Monte Carlo methods for estimating the diagonal of a real symmetric matrix”
Eric Hallman, Ilse Ipsen and Arvind Saibaba · 2023
Later among the works it cites.
“Smoothing methods for automatic differentiation across conditional branches”
Justin Kreikemeyer and Philipp Andelfinger · 2023
Later among the works it cites.
“Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training”
Hong Liu et al · 2023
Later among the works it cites.
“Probabilistic Machine Learning: Advanced Topics”
Kevin. Murphy · 2023
Later among the works it cites.
“Optimisation in Neurosymbolic Learning Systems”, 2024
Emile van Krieken · 2024
Closest in time.