Fetching the paper…
Reading the bibliography…
Sign-based optimization methods have become popular in machine learning due to their favorable communication cost in distributed optimization and their surprisingly good performance in neural network training.
Analysis of gradient clipping and adaptive scaling with a relaxed smoothness condition
Zhang, J., He, T., Sra, S., and Jadbabaie, A · 1905
Earlier work this paper cites.
Why Adam beats SGD for attention models
Zhang, J., Karimireddy, S. P., Veit, A., Kim, S., Reddi, S. J., Kumar, S., and Sra, S · 1912
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second-order methods. i-voc. 1988 connectionist models summer school, 1988
Becker, S. and LeCun, Y · 1988
Earlier work this paper cites.
Computing the norm ‖ A ‖ ∞ , 1 \|A\|_{\infty,1} is NP-hard
Rohn, J · 2000
Earlier work this paper cites.
SciPy: Open source scientific tools for Python, 2001
Jones, E., Oliphant, T., Peterson, P., et al · 2001
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
An almost-linear-time algorithm for approximate max flow in undirected graphs, and its multicommodity generalizations
Kelner, J. A., Lee, Y. T., Orecchia, L., and Sidford, A · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Earlier work this paper cites.
Stochastic spectral descent for discrete graphical models
Carlson, D., Hsieh, Y.-P., Collins, E., Carin, L., and Cevher, V · 2015
Earlier work this paper cites.
Beyond convexity: Stochastic quasi-convex optimization
Hazan, E., Levy, K., and Shalev-Shwartz, S · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Cited alongside, same era.
The power of normalization: Faster evasion of saddle points
Levy, K. Y · 2016
Cited alongside, same era.
Entropy-SGD: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J. T., Sagun, L., and Zecchina, R · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., and Wanderman-Milne, S · 2018
Later among the works it cites.
The full spectrum of deep net Hessians at scale: Dynamics with sample size
Papyan, V · 2018
Later among the works it cites.
Adolphs, L., Kohler, J., and Lucchi, A · 2019
Later among the works it cites.
SignSGD with majority vote is communication efficient and fault tolerant
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A · 2019
Later among the works it cites.
An investigation into neural net optimization via Hessian eigenvalue density
Ghorbani, B., Krishnan, S., and Xiao, Y · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Cited alongside, same era.
Block-normalized gradient method: An empirical study for training deep neural network
Yu, A. W., Huang, L., Lin, Q., Salakhutdinov, R., and Carbonell, J · 2017
Cited alongside, same era.
Dissecting Adam: The sign, magnitude and variance of stochastic gradients
Balles, L. and Hennig, P · 2018
Cited alongside, same era.
SignSGD: Compressed optimisation for non-convex problems
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A · 2018
Cited alongside, same era.
Later among the works it cites.
Stochastic gradient methods with layer-wise adaptive moments for training of deep networks
Ginsburg, B., Castonguay, P., Hrinchuk, O., Kuchaiev, O., Lavrukhin, V., Leary, R., Li, J., Nguyen, H., and Cohen, J. M · 2019
Later among the works it cites.
Error feedback fixes SignSGD and other gradient compression schemes
Karimireddy, S. P., Rebjock, Q., Stich, S., and Jaggi, M · 2019
Later among the works it cites.
Hessian based analysis of SGD for deep nets: Dynamics and generalization
Li, X., Gu, Q., Zhou, Y., Chen, T., and Banerjee, A · 2019
Later among the works it cites.
On stochastic sign descent methods
Safaryan, M. and Richtárik, P · 2019
Later among the works it cites.