A universally optimal multistage accelerated stochastic gradient method
Aybat, N. S., Fallah, A., Gurbuzbalaban, M., and Ozdaglar, A · 2019
Later among the works it cites.
Deep Frank-Wolfe for neural network optimization
Berrada, L., Zisserman, A., and Kumar, M · 2019
Later among the works it cites.
Autoaugment: Learning augmentation strategies from data
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V · 2019
Later among the works it cites.
Stochastic algorithms with geometric step decay converge linearly on sharp functions
Original
Davis, D., Drusvyatskiy, D., and Charisopoulos, V · 2019
Later among the works it cites.
The step decay schedule: A near optimal, geometrically decaying learning rate procedure for least squares
Ge, R., Kakade, S. M., Kidambi, R., and Netrapalli, P · 2019
Later among the works it cites.
Stochastic gradient methods with layer-wise adaptive moments for training of deep networks
Original
Ginsburg, B., Castonguay, P., Hrinchuk, O., Kuchaiev, O., Lavrukhin, V., Leary, R., Li, J., Nguyen, H., Zhang, Y., and Cohen, J. M · 2019
Later among the works it cites.
Bag of tricks for image classification with convolutional neural networks
He, T., Zhang, Z., Zhang, H., Zhang, Z., Xie, J., and Li, M · 2019
Later among the works it cites.
A generic acceleration framework for stochastic composite optimization
Kulunchakov, A. and Mairal, J · 2019
Later among the works it cites.
Attention network robustification for person ReID
Original
Lawen, H., Ben-Cohen, A., Protter, M., Friedman, I., and Zelnik-Manor, L · 2019
Later among the works it cites.
On the convergence of stochastic gradient descent with adaptive stepsizes
Li, X. and Orabona, F · 2019
Later among the works it cites.
A modern introduction to online learning
Original
Orabona, F · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes
Ward, R., Wu, X., and Bottou, L · 2019
Later among the works it cites.
Stagewise training accelerates convergence of testing error over sgd
Yuan, Z., Yan, Y., Jin, R., and Yang, T · 2019
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Closest in time.
Bootstrap your own latent - a new approach to self-supervised learning
Grill, J., Strub, F., Altché, F., Tallec, C., Richemond, P., Buchatskaya, E., Doersch, C., Avila P., B., Guo, Z., Gheshlaghi Azar, M., Piot, B., kavukcuoglu, K., Munos, R., and Valko, M · 2020
Closest in time.
Better theory for SGD in the nonconvex world
Original
Khaled, A. and Richtárik, P · 2020
Closest in time.
A high probability analysis of adaptive SGD with momentum
Li, X. and Orabona, F · 2020
Closest in time.
Page: A simple and optimal probabilistic gradient estimator for nonconvex optimization
Original
Li, Z., Bao, H., Zhang, X., and Richtárik, P · 2020
Closest in time.
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes
Ward, R., Wu, X., and Bottou, L · 2020
Closest in time.
Graph structure of neural networks
You, J., Leskovec, J., He, K., and Xie, S · 2020
Closest in time.
Exploring self-attention for image recognition
Zhao, H., Jia, J., and Koltun, V · 2020
Closest in time.
From low probability to high confidence in stochastic convex optimization
Davis, D., Drusvyatskiy, D., Xiao, L., and Zhang, J · 2021
Closest in time.