Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P., Ying, C., and Le, Q. V · 2018
Later among the works it cites.
Global convergence of Langevin dynamics based algorithms for nonconvex optimization
Xu, P., Chen, J., Zou, D., and Gu, Q · 2018
Later among the works it cites.
Stochastic fractional Hamiltonian Monte Carlo
Ye, N. and Zhu, Z · 2018
Later among the works it cites.
Stochastic variance-reduced Hamilton Monte Carlo methods
Zou, D., Xu, P., and Gu, Q · 2018
Later among the works it cites.
The tamed unadjusted Langevin algorithm
Brosse, N., Durmus, A., Moulines, É., and Sabanis, S · 2019
Later among the works it cites.
Stationary states for underdamped anharmonic oscillators driven by Cauchy noise
Original
Capała, K. and Dybiec, B · 2019
Later among the works it cites.
On the diffusion approximation of nonconvex stochastic gradient descent
Hu, W., Li, C. J., Li, L., and Liu, J.-G · 2019
Later among the works it cites.
Kinetic energy choice in Hamiltonian/hybrid Monte Carlo
Livingstone, S., Faulkner, M. F., and Roberts, G. O · 2019
Later among the works it cites.
Heavy-tailed universality predicts trends in test accuracies for very large pre-trained deep neural networks
Original
Martin, C. H. and Mahoney, M. W · 2019
Later among the works it cites.
First exit time analysis of stochastic gradient descent under heavy-tailed gradient noise
Nguyen, T. H., Şimşekli, U., Gurbuzbalaban, M., and Richard, G · 2019
Later among the works it cites.
Non-Asymptotic Analysis of Fractional Langevin Monte Carlo for Non-Convex Optimization
Nguyen, T. H., Şimşekli, U., and Richard, G · 2019
Later among the works it cites.
Non-Gaussianity of stochastic gradient noise
Original
Panigrahi, A., Somani, R., Goyal, N., and Netrapalli, P · 2019
Later among the works it cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhu, Z., Wu, J., Yu, B., Wu, L., and Ma, J · 2019
Later among the works it cites.
Stochastic gradient Hamiltonian Monte Carlo methods with recursive variance reduction
Zou, D., Xu, P., and Gu, Q · 2019
Later among the works it cites.
Hausdorff dimension, stochastic differential equations, and generalization in neural networks
Original
Şimşekli, U., Sener, O., Deligiannidis, G., and Erdogdu, M. A · 2020
Closest in time.
On sampling from a log-concave density using kinetic Langevin diffusions
Dalalyan, A. S. and Riou-Durand, L · 2020
Closest in time.
The heavy-tail phenomenon in SGD
Original
Gürbüzbalaban, M., Şimşekli, U., and Zhu, L · 2020
Closest in time.
Multiplicative noise and heavy tails in stochastic optimization
Original
Hodgkinson, L. and Mahoney, M. W · 2020
Closest in time.