Non-convex learning via stochastic gradient Langevin dynamics: a nonasymptotic analysis
M. Raginsky, A. Rakhlin, and M. Telgarsky · 2017
Later among the works it cites.
Empirical analysis of the hessian of over-parametrized neural networks
Original
Levent Sagun, Utku Evci, V. Uğur Güney, Yann Dauphin, and Léon Bottou · 2017
Later among the works it cites.
Fractional Langevin Monte Carlo: Exploring L \ \backslash ’ { \{ e } \} vy Driven Stochastic Differential Equations for Markov Chain Monte Carlo
Original
Umut Şimşekli · 2017
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Original
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Later among the works it cites.
Comparing dynamics: Deep neural networks versus glassy systems
Marco Baity-Jesi, Levent Sagun, Mario Geiger, Stefano Spigler, Gerard Ben Arous, Chiara Cammarota, Yann LeCun, Matthieu Wyart, and Giulio Biroli · 2018
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
P. Chaudhari and S. Soatto · 2018
Later among the works it cites.
Escaping saddles with stochastic gradients
H. Daneshmand, J. Kohler, A. Lucchi, and T. Hofmann · 2018
Later among the works it cites.
Global non-convex optimization with discretized diffusions
M. A. Erdogdu, L. Mackey, and O. Shamir · 2018
Later among the works it cites.
Rethinking learning rate schedules for stochastic optimization
Rong Ge, Sham M Kakade, Rahul Kidambi, and Praneeth Netrapalli · 2018
Later among the works it cites.
The jamming transition as a paradigm to understand the loss landscape of deep neural networks
Original
Mario Geiger, Stefano Spigler, Stéphane d’Ascoli, Levent Sagun, Marco Baity-Jesi, Giulio Biroli, and Matthieu Wyart · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot-Guillarmod, Franck Gabriel, and Clement Hongler · 2018
Later among the works it cites.
Deep neural networks as gaussian processes
Jae Hoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
Gradient estimates and ergodicity for sdes driven by multiplicative lévy noises via coupling
Original
Mingjie Liang and Jian Wang · 2018
Later among the works it cites.
Revisiting small batch training for deep neural networks
Original
Dominic Masters and Carlo Luschi · 2018
Later among the works it cites.
On the convergence of gradient-like flows with noisy gradient input
Panayotis Mertikopoulos and Mathias Staudigl · 2018
Later among the works it cites.
The full spectrum of deep net Hessians at scale: Dynamics with sample size
Original
Vardan Papyan · 2018
Later among the works it cites.
Alpha-stable low-rank plus residual decomposition for speech enhancement
U. Şimşekli, H. Erdoğan, S. Leglaive, A. Liutkus, R. Badeau, and G. Richard · 2018
Later among the works it cites.
Local optimality and generalization guarantees for the langevin algorithm via empirical metastability
B. Tzen, T. Liang, and M. Raginsky · 2018
Later among the works it cites.
A walk with sgd
Original
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio · 2018
Later among the works it cites.
Global convergence of langevin dynamics based algorithms for nonconvex optimization
Pan Xu, Jinghui Chen, Difan Zou, and Quanquan Gu · 2018
Later among the works it cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from minima and regularization effects
Original
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2018
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
U. Şimşekli, L. Sagun, and Gürbüzbalaban · 2019
Closest in time.
Stochastic gradient descent for nonconvex learning without bounded gradient assumptions
Original
Yunwen Lei, Ting Hu, and Ke Tang · 2019
Closest in time.
Traditional and heavy-tailed self regularization in neural network models
Original
Charles H Martin and Michael W Mahoney · 2019
Closest in time.
Non-gaussianity of stochastic gradient noise
Original
Abhishek Panigrahi, Raghav Somani, Navin Goyal, and Praneeth Netrapalli · 2019
Closest in time.
The effect of network width on stochastic gradient descent and generalization: an empirical study
Original
Daniel S Park, Jascha Sohl-Dickstein, Quoc V Le, and Samuel L Smith · 2019
Closest in time.
Fluctuation-dissipation relations for stochastic gradient descent
S. Yaida · 2019
Closest in time.
A hitting time analysis of stochastic gradient langevin dynamics
Y. Zhang, P. Liang, and M. Charikar · 2022
Closest in time.