Fetching the paper…
Reading the bibliography…
Over the past few years, deep learning has risen to the foreground as a topic of massive interest, mainly as a result of successes obtained in solving large-scale image processing tasks.
Note on the Derivatives with Respect to a Parameter of the Solutions of a System of Differential Equations
Thomas H. Grönwall · 1919
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Foundations of differential geometry. Vol. I
Shoshichi Kobayashi and Katsumi Nomizu · 1963
Earlier work this paper cites.
The representation of the cumulative rounding error of an algorithm as a taylor expansion of the local rounding errors
Seppo Linnainmaa · 1970
Earlier work this paper cites.
Generalized disks of contractivity for explicit and implicit Runge-Kutta methods
Germund Dahlquist · 1979
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
John J. Hopfield · 1982
Earlier work this paper cites.
Generalized canonical transformations for time-dependent systems
Manuel Asorey, José F. Cariñena, and Luis A. Ibort · 1983
Earlier work this paper cites.
Mathematical Theory of Optimal Processes
L S Pontryagin · 1987
Earlier work this paper cites.
A theoretical framework for back-propagation
Yann LeCun · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Yann LeCun, Bernhard Boser, John S. Denker, Donnie Henderson, Richard E. Howard, Wayne Hubbard, and Lawrence D. Jackel · 1989
Earlier work this paper cites.
A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines
Michael F. Hutchinson · 1990
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
Equivariant adaptive source separation
Jean-François Cardoso and Beate Hvam Laheld · 1992
Earlier work this paper cites.
Solving Ordinary Differential Equations I
Ernst Hairer, Syvert P. Nørsett, and Gerhard Wanner · 1993
Earlier work this paper cites.
Convex functions and optimization methods on Riemannian manifolds
Constantin Udrişte · 1994
Earlier work this paper cites.
A new learning algorithm for blind signal separation
Shun-Ichi Amari, Andrzej Cichocki, and Howard Hua Yang · 1996
Earlier work this paper cites.
Regularization of Inverse Problems
Heinz Werner Engl, Martin Hanke, and Andreas Neubauer · 1996
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
An overview of ERS-SAR interferometry
Fabio Rocca, Claudio Maria Prato, and Alessandro Ferretti · 1997
Earlier work this paper cites.
Natural gradient descent for training multi-layer perceptrons
Howard Hua Yang and Shun-ichi Amari · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Why natural gradient?
Shun-Ichi Amari and Scott C. Douglas · 1998
Earlier work this paper cites.
Convolutional networks for images, speech, and time series, 1998
Yann LeCun and Yoshua Bengio · 1998
Earlier work this paper cites.
Geometric integration using discrete gradients
Robert I. McLachlan, G. Reinout W. Quispel, and Nicolas Robidoux · 1999
Earlier work this paper cites.
Trust-Region Methods
Andrew R. Conn, Nicholas I. M. Gould, and Philippe L. Toint · 2000
Earlier work this paper cites.
Runge-Kutta methods in optimal control and the transformed adjoint system
William W. Hager · 2000
Earlier work this paper cites.
Independent component analysis: algorithms and applications
Aapo Hyvärinen and Erkki Oja · 2000
Earlier work this paper cites.
Lie-group methods
Arieh Iserles, Hans Z. Munthe-Kaas, Syvert P. Nørsett, and Antonella Zanna · 2000
Earlier work this paper cites.
Regularization of Inverse Problems, 2020
Christian Clason · 2001
Earlier work this paper cites.
Conformal Hamiltonian systems
Robert McLachlan and Matthew Perlmutter · 2001
Earlier work this paper cites.
Splitting methods
Robert I. McLachlan and G. Reinout W. Quispel · 2002
Earlier work this paper cites.
Motion capture database
Carnegie Mellon University Graphics Lab · 2003
Earlier work this paper cites.
Neural learning by geometric integration of reduced ‘rigid-body’ equations
Elena Celledoni and Simone Fiori · 2004
Earlier work this paper cites.
Riemannian geometry
Sylvestre Gallot, Dominique Hulin, and Jacques Lafontaine · 2004
Earlier work this paper cites.
Feature selection, L 1 L_{1} vs. L 2 L_{2} regularization, and rotational invariance
Andrew Y. Ng · 2004
Earlier work this paper cites.
Camino: Open-Source Diffusion-MRI Reconstruction and Processing
Phil Cook, Yu Bai, Shahrum Nedjati-Gilani, Kiran Seunarine, Matt Hall, Geoffrey Parker, and Daniel Alexander · 2006
Earlier work this paper cites.
Geometric numerical integration: structure-preserving algorithms for ordinary differential equations
Ernst Hairer, Christian Lubich, and Gerhard Wanner · 2006
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Measure theory
Vladimir I. Bogachev · 2007
Earlier work this paper cites.
Optimization algorithms on matrix manifolds
Pierre-Antoine Absil, Robert Mahony, and Rodolphe Sepulchre · 2008
Earlier work this paper cites.
Gradient flows: in metric spaces and in the space of probability measures
Luigi Ambrosio, Nicola Gigli, and Giuseppe Savaré · 2008
Earlier work this paper cites.
Efficient Projections onto the l 1 l_{1} -ball for Learning in High dimensions
John Duchi, Shai Shalev-Shwartz, Yoram Singer, and Tushar Chandra · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Cited alongside, same era.
Sparse feature learning for deep belief networks
Marc’Aurelio Ranzato, Y-Lan Boureau, and Yann Le Cun · 2009
Cited alongside, same era.
Solving ordinary differential equations. II
Ernst Hairer and Gerhard Wanner · 2010
Cited alongside, same era.
Stacked denoising autoencoders: learning useful representations in a deep network with a local denoising criterion
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol · 2010
Cited alongside, same era.
log det A = tr log A
Christopher S. Withers and Saralees Nadarajah · 2010
Cited alongside, same era.
ImageNet Classification with Deep Convolutional Neural Networks, 2012
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Learning SO(3) Equivariant Representations with Spherical CNNs
Carlos Esteves, Christine Allen-Blanchette, Ameesh Makadia, and Kostas Daniilidis · 2018
Later among the works it cites.
i-RevNet: Deep invertible networks
Jörn-Henrik Jacobsen, Arnold W.M. Smeulders, and Edouard Oyallon · 2018
Later among the works it cites.
Glow: Generative flow with invertible 1x1 convolutions
Diederik P. Kingma and Prafulla Dhariwal · 2018
Later among the works it cites.
Clebsch–Gordan Nets: a Fully Fourier Space Spherical Convolutional Neural Network, 2018
Risi Kondor, Zhen Lin, and Shubhendu Trivedi · 2018
Later among the works it cites.
Risi Kondor and Shubhendu Trivedi · 2018
Later among the works it cites.
Maximum principle based algorithms for deep learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2013
Cited alongside, same era.
An introduction to Lie group integrators—basics, new developments and applications
Elena Celledoni, Håkon Marthinsen, and Brynjulf Owren · 2014
Cited alongside, same era.
NICE: Non-linear independent components estimation
Laurent Dinh, David Krueger, and Yoshua Bengio · 2014
Cited alongside, same era.
Inverse Problems - Tikhonov Theory and Algorithms
Kazufumi Ito and Bangti Jin · 2014
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Qianxiao Li, Long Chen, Cheng Tai, and Weinan E · 2018
Later among the works it cites.
Beyond finite layer neural networks: Bridging deep architectures and numerical differential equations
Yiping Lu, Aoxiao Zhong, Quanzheng Li, and Bin Dong · 2018
Later among the works it cites.
Chris J Maddison, Daniel Paulin, Yee Whye Teh, Brendan O’Donoghue, and Arnaud Doucet · 2018
Later among the works it cites.
On the convergence of Adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
Tensor field networks: Rotation- and translation-equivariant neural networks for 3D point clouds
Nathaniel Thomas, Tess Smidt, Steven Kearnes, Lusann Yang, Li Li, Kai Kohlhoff, and Patrick Riley · 2018
Later among the works it cites.
Deep limits of residual neural networks
Matthew Thorpe and Yves van Gennip · 2018
Later among the works it cites.
Deep Image Prior
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2018
Later among the works it cites.
Learning Steerable Filters for Rotation Equivariant CNNs
Maurice Weiler, Fred A. Hamprecht, and Martin Storath · 2018
Later among the works it cites.
Universal approximations of invariant maps by neural networks
Dmitry Yarotsky · 2018
Later among the works it cites.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
Solving inverse problems using data-driven models
Simon Arridge, Peter Maass, Ozan Öktem, and Carola-Bibiane Schönlieb · 2019
Later among the works it cites.
Riemannian adaptive optimization methods
Gary Bécigneul and Octavian-Eugen Ganea · 2019
Later among the works it cites.
Invertible residual networks
Jens Behrmann, Will Grathwohl, Ricky T. Q. Chen, David Duvenaud, and Joern-Henrik Jacobsen · 2019
Later among the works it cites.
Deep learning as optimal control problems: models and numerical methods
Martin Benning, Elena Celledoni, Matthias J. Ehrhardt, Brynjulf Owren, and Carola-Bibiane Schönlieb · 2019
Later among the works it cites.
Optimal Approximation with Sparsely Connected Deep Neural Networks
Helmut Bölcskei, Philipp Grohs, Gitta Kutyniok, and Philipp Petersen · 2019
Later among the works it cites.
Course on Optimal Control, 2019
J. Frédéric Bonnans · 2019
Later among the works it cites.
Residual flows for invertible generative modeling
Tian Qi Chen, Jens Behrmann, David Duvenaud, and Jörn-Henrik Jacobsen · 2019
Later among the works it cites.
A General Theory of Equivariant CNNs on Homogeneous Spaces
Taco Cohen, Mario Geiger, and Maurice Weiler · 2019
Later among the works it cites.
Augmented Neural ODEs
Emilien Dupont, Arnaud Doucet, and Yee Whye Teh · 2019
Later among the works it cites.
Neural spline flows
Conor Durkan, Artur Bekasov, Iain Murray, and George Papamakarios · 2019
Later among the works it cites.
Conformal symplectic and relativistic optimization
Guilherme França, Jeremias Sulam, Daniel P. Robinson, and René Vidal · 2019
Later among the works it cites.
ANODE: Unconditionally accurate memory-efficient gradients for neural ODEs
Amir Gholami, Kurt Keutzer, and George Biros · 2019
Later among the works it cites.
Emerging convolutions for generative normalizing flows
Emiel Hoogeboom, Rianne Van Den Berg, and Max Welling · 2019
Later among the works it cites.
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila · 2019
Later among the works it cites.
Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley Transform
Jun Li, Fuxin Li, and Sinisa Todorovic · 2019
Later among the works it cites.
Port-Hamiltonian approach to neural network training
Stefano Massaroli, Michael Poli, Federico Califano, Angela Faragasso, Jinkyoo Park, Atsushi Yamashita, and Hajime Asama · 2019
Later among the works it cites.
Hamiltonian descent for composite objectives
Brendan O’Donoghue and Chris J. Maddison · 2019
Later among the works it cites.
Equivalence of approximation by convolutional neural networks and fully-connected networks
Philipp Petersen and Felix Voigtlaender · 2019
Later among the works it cites.
Invert to learn to invert
Patrick Putzky and Max Welling · 2019
Later among the works it cites.
Deep Neural Networks Motivated by Partial Differential Equations
Lars Ruthotto and Eldad Haber · 2019
Later among the works it cites.
Fast convergence of natural gradient descent for over-parameterized neural networks
Guodong Zhang, James Martens, and Roger B Grosse · 2019
Later among the works it cites.
On the invertibility of invertible neural networks, 2020
Jens Behrmann, Paul Vicol, Kuan-Chieh Wang, Roger B. Grosse, and Jörn-Henrik Jacobsen · 2020
Closest in time.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lénaïc Chizat and Francis Bach · 2020
Closest in time.
iUNets: Fully invertible U-Nets with learnable up-and downsampling
Christian Etmann, Rihuan Ke, and Carola-Bibiane Schönlieb · 2020
Closest in time.
Efficient Riemannian optimization on the Stiefel manifold via the Cayley transform
Sinisa Todorovic Jun Li, Li Fuxin · 2020
Closest in time.
Analysis of the BFGS Method with Errors
Yuchen Xie, Richard H. Byrd, and Jorge Nocedal · 2020
Closest in time.
Forward Stability of ResNet and Its Variants
Linan Zhang and Hayden Schaeffer · 2020
Closest in time.