Understand
Effective training of deep neural networks suffers from two main issues.
- The first is that the parameter spaces of these models exhibit pathological curvature.
- Recent methods address this problem by using adaptive preconditioning for Stochastic Gradient Descent (SGD).
- These methods improve convergence by adapting to the local geometry of parameter space.
Built on
Nothing clear enough to list yet.
Similar
Nothing clear enough to list yet.
Then
Nothing clear enough to list yet.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…