Fetching the paper…
Reading the bibliography…
We prove that Riemannian contraction in a supervised learning setting implies generalization.
Gradient methods for the minimisation of functionals
Boris Polyak · 1963
Earlier work this paper cites.
Solution of the matrix equation ax + xb = c [f4]
R. H. Bartels and G. W. Stewart · 1972
Earlier work this paper cites.
Applied nonlinear control , volume 199
Jean-Jacques E Slotine, Weiping Li, et al · 1991
Earlier work this paper cites.
Principal components, minor components, and linear neural networks
Erkki Oja · 1992
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
On contraction analysis for non-linear systems
Winfried Lohmiller and Jean-Jacques E Slotine · 1998
Earlier work this paper cites.
An overview of statistical learning theory
Vladimir N Vapnik · 1999
Earlier work this paper cites.
Modularity, evolution, and the binding problem: a view from stability theory
J.-J.E. Slotine and W. Lohmiller · 2001
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Kernel methods for pattern analysis
John Shawe-Taylor, Nello Cristianini, et al · 2004
Earlier work this paper cites.
Stability of randomized learning algorithms
Andre Elisseeff, Theodoros Evgeniou, Massimiliano Pontil, and Leslie Pack Kaelbing · 2005
Earlier work this paper cites.
On partial contraction analysis for coupled nonlinear oscillators
Wei Wang and Jean-Jacques E Slotine · 2005
Earlier work this paper cites.
Learning theory: stability is sufficient for generalization and necessary and sufficient for consistency of empirical risk minimization
Sayan Mukherjee, Partha Niyogi, Tomaso Poggio, and Ryan Rifkin · 2006
Earlier work this paper cites.
Stochastic convex optimization
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2009
Earlier work this paper cites.
Learnability, stability and uniform convergence
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2010
Earlier work this paper cites.
A contraction theory approach to singularly perturbed systems
Domitilla Del Vecchio and Jean-Jacques E Slotine · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Cited alongside, same era.
Stochastic contraction in riemannian metrics
Quang-Cuong Pham and Jean-Jacques Slotine · 2013
Cited alongside, same era.
Contraction analysis of nonlinear random dynamical systems
Nicolas Tabareau and Jean-Jacques Slotine · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Continuous-time limit of stochastic gradient descent revisited
Stephan Mandt, Matthew D Hoffman, David M Blei, et al · 2015
Cited alongside, same era.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Later among the works it cites.
Fast convergence of natural gradient descent for overparameterized neural networks
Guodong Zhang, James Martens, and Roger Grosse · 2019
Later among the works it cites.
A continuous-time analysis of distributed stochastic gradient
Nicholas M Boffi and Jean-Jacques E Slotine · 2020
Later among the works it cites.
Sharper bounds for uniformly stable algorithms
Olivier Bousquet, Yegor Klochkov, and Nikita Zhivotovskiy · 2020
Later among the works it cites.
Stanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani, Daniel M Roy, and Surya Ganguli · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent, 2016
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2016
Cited alongside, same era.
Stability and generalization of learning algorithms that converge to global optima
Zachary Charles and Dimitris Papailiopoulos · 2018
Cited alongside, same era.
Stability and convergence trade-off of iterative optimization algorithms
Yuansi Chen, Chi Jin, and Bin Yu · 2018
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Later among the works it cites.
Instability, computational efficiency and statistical accuracy
Nhat Ho, Koulik Khamaru, Raaz Dwivedi, Martin J Wainwright, Michael I Jordan, and Bin Yu · 2020
Later among the works it cites.
On the choice of metric in gradient-based theories of brain function
Simone Carlo Surace, Jean-Pascal Pfister, Wulfram Gerstner, and Johanni Brea · 2020
Later among the works it cites.
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J. van der Walt, Matthew Brett, Joshua Wilson, K. Jarrod Millman, Nikolay Mayorov, Andrew R. J. Nelson, Eric Jones, Robert Kern, Eric Larson, C J Carey, İlhan Polat, Yu Feng, Eric W. Moore, Jake VanderPlas, Denis Laxalde, Josef Perktold, Robert Cimrman, Ian Henriksen, E. A. Quintero, Charles R. Harris, Anne M. Archibald, Antônio H. Ribeiro, Fabian Pedregosa, Paul van Mulbregt, and SciPy 1.0 Contributors · 2020
Later among the works it cites.
Beyond convexity—contraction and global convergence of gradient descent
Patrick M Wensing and Jean-Jacques Slotine · 2020
Later among the works it cites.
Algorithmic instabilities of accelerated gradient descent
Amit Attia and Tomer Koren · 2021
Later among the works it cites.
Regret bounds for adaptive nonlinear control
Nicholas M Boffi, Stephen Tu, and Jean-Jacques E Slotine · 2021
Later among the works it cites.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Later among the works it cites.
Recursive construction of stable assemblies of recurrent neural networks
Leo Kozachkov, Michaela Ennis, and Jean-Jacques Slotine · 2021
Later among the works it cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks, 2021
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2021
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Later among the works it cites.